16 September 2026 · LLM × Safety Review × AE

Can an LLM find adverse events that human reviewers miss?

A 30-minute pack on LLMs as second readers, sensitivity versus precision, disagreement-driven adjudication, controlled terminology, and the path from narrative evidence to ADAE and TFLs.

DifficultyC1
Time30 minutes
Main sourceJAMA · 3 Sep
OutputAE review workflow

Why this reading

AI and human reviewers do not necessarily make the same mistakes.

A fresh open-access JAMA Network Open study compares an LLM with reviewer consensus for AE identification in encounter notes from four randomized immunotherapy trials. For SP work, it connects narrative source evidence with safety review, terminology, structured AE data and downstream analysis.

The most useful design may be human review + machine review → disagreement set → targeted adjudication, rather than autonomous replacement.

Reading order

Your 30-minute plan.

0–3 minPreview

Predict whether the model will favor sensitivity or precision.

3–14 minJAMA

Study reference standard, performance metrics and disagreement cases.

14–20 minTerminology

Connect note-level concepts to standardized representation.

20–25 minWorkflow

Place extraction and validation inside machine-readable clinical pipelines.

25–30 minOutput

Design disagreement-driven AE review and escalation.

Open-access sources

Primary evidence plus standards context.

Brief background

Detection is upstream of coding, SDTM, ADaM and the final TFL.

Narrative notes can contain safety evidence that is difficult to capture with simple structured fields. An LLM can act as a second reader, but detection is not equivalent to coding or analysis.

Sensitivity and precision create different operational trade-offs. In a screening layer, extra false positives may mean more review; false negatives may mean missed safety information. The consequences are asymmetric.

A defensible evidence chain is note ID → supporting text → extracted AE → human decision → coded term → SDTM AE → ADAE → TFL.

Key vocabulary

Fifteen terms for AI-assisted safety review.

Term中文Meaning
adverse event (AE)不良事件Any unfavorable medical occurrence after an intervention; causality is not required.
reviewer consensus审阅者共识A reference decision formed from multiple human judgments.
sensitivity灵敏度Proportion of true target events successfully identified.
precision精确率Proportion of identified events that are actually correct.
false positive假阳性An identified event not supported by the reference assessment.
false negative假阴性A real event that the system fails to identify.
clinical note临床记录Narrative documentation written during care or trial encounters.
information extraction信息抽取Converting facts in free text into structured concepts.
controlled terminology受控术语Standardized permitted terms used consistently.
MedDRA codingMedDRA 编码Mapping an AE concept to MedDRA.
human-in-the-loop人在回路Humans review or resolve AI-generated results.
adjudication裁定Formal expert resolution of disagreement or uncertainty.
traceability可追溯性Connecting a derived value to source evidence and transformation history.
signal detection信号检测Identifying patterns that may indicate a safety issue.
error asymmetry错误不对称性Different error types have different downstream consequences.

Useful phrases

Language for safety and QC discussions.

  1. The model should assist detection rather than become the source of truth.
  2. High sensitivity can be useful when missed events are costly.
  3. False positives create review burden but false negatives may create safety risk.
  4. The extracted event should remain linked to the supporting note.
  5. Human consensus is useful but is not automatically infallible.
  6. Disagreement cases deserve targeted adjudication.
  7. Standard terminology is required before narrative concepts become analysis categories.
  8. The optimal threshold depends on the downstream decision.
  9. Automation changes the review workload rather than eliminating review.
  10. Performance metrics should be interpreted in the context of the task.

Comprehension

Five questions.

  1. Why is reviewer consensus useful but not automatically infallible?
  2. What is the operational difference between sensitivity and precision?
  3. Why may false negatives be especially costly in safety screening?
  4. Why is note-level AE detection not equivalent to SDTM AE or ADAE creation?
  5. How can human–AI disagreement become a QC signal?

Retelling

Say it three times.

  • 30 seconds · LLM second reader versus human consensus.
  • 45 seconds · Sensitivity, precision, false positives and false negatives.
  • 60 seconds · Clinical note → AE evidence → coding → SDTM → ADAE → TFL.

5-minute output task

Design an AI-assisted AE review workflow.

  1. Minute 1: Define what AI may detect or suggest.
  2. Minutes 2–3: Define actions for agreement, AI-only detection and human-only detection.
  3. Minute 4: Define the evidence package and downstream linkage.
  4. Minute 5: Explain where human judgment remains mandatory.

One sentence to keep

In clinical safety review, the most useful AI may be a second reader whose disagreements with humans reveal exactly where additional judgment is needed.