PHQ-9 and GAD-7: what clinicians actually watch for
PHQ-9 and GAD-7 are the two screens that show up in nearly every behavioral-health workflow. The thing nobody teaches outside of training: the score is the least interesting part. Here is what to read instead.

PHQ-9 and GAD-7 are the two screens that show up in nearly every outpatient behavioral-health workflow. They are short, validated, free to use, and they fit on a single sheet. They are also the two instruments most often misread, because the score is the least interesting thing they produce.
Here is what to read instead.
The score is a starting point, not a verdict
PHQ-9 is a nine-item self-report that returns a total between 0 and 27. The validated severity bands are:
- 0 to 4 Minimal or none
- 5 to 9 Mild
- 10 to 14 Moderate
- 15 to 19 Moderately severe
- 20 to 27 Severe
GAD-7 is seven items, 0 to 21:
- 0 to 4 Minimal
- 5 to 9 Mild
- 10 to 14 Moderate
- 15 to 21 Severe
Two patients can land at PHQ-9 = 12 ("moderate") for very different reasons. One is sleeping poorly, eating poorly, and feeling low energy. All somatic items, no suicidality, no anhedonia. The other has flat affect, no interest in things they used to love, and a 1 on item 9. Same total. Very different clinical picture.
The score is a triage signal. The item pattern is the diagnostic signal.
Item 9 is the line in the sand
PHQ-9 item 9 ("Thoughts that you would be better off dead or of hurting yourself in some way") is the only item on either instrument that requires action regardless of the rest of the score. A patient with a total PHQ-9 of 6 and a 1 on item 9 is not a "mild" patient. They are a patient who needs a safety check today.
This is also where automated systems get themselves in the most trouble. Any tool that administers PHQ-9 and does not escalate immediately on a non-zero item 9, whether by surfacing crisis resources, opening a clinician notification, or both, is misusing the instrument. The 988 Suicide and Crisis Lifeline (call or text 988 in the US) and 911 for immediate danger should be visible as plain readable text the moment that item is non-zero, not buried behind an icon and not hidden behind a confirmation step. We wrote a longer post on item-9 escalation design for the full breakdown of how this should be implemented.
If your workflow has ever surfaced a "we noticed something concerning" message instead of the actual phone number, that is a failure mode worth fixing this week.
What movement on the score actually means
A single PHQ-9 score is a snapshot. The clinically meaningful thing is change over time, ideally with the same instrument administered at consistent intervals. The literature treats a five-point change in PHQ-9 as the threshold for a minimal clinically important difference; smaller swings are within the noise of the instrument.
This is why a single-administration mood-tracking app misses the point. PHQ-9 once at intake gives you a number. PHQ-9 at days 1, 7, and 14 of a treatment arc gives you a trajectory, and trajectory is what therapy and pharmacotherapy actually move. A patient going from 18 to 14 over two weeks is responding. A patient going from 18 to 17 is not, regardless of which side of "moderately severe" they sit on.
GAD-7 is similar but quieter. The instrument is sensitive to changes of about four points; below that, treat shifts as informational rather than as evidence of response.
Don't reword the items
This is the rule that gets broken the most often. The validated wording of PHQ-9 and GAD-7 is the validated wording. "Little interest or pleasure in doing things" must stay "little interest or pleasure in doing things." "Feeling down, depressed, or hopeless" must stay "feeling down, depressed, or hopeless." The four response options (Not at all, Several days, More than half the days, Nearly every day) must stay in that exact order with that exact text.
This sounds like pedantry until you remember why: the validation studies were run with that exact text. Change one item to "Have you been struggling to find joy lately?" and you have invented a new instrument with no published reliability, no published sensitivity, no published specificity. Nobody can tell you what a score of 12 means on your new version.
The same principle applies to translations. Validated translations exist for both instruments in many languages; use those. Do not run the English text through an LLM and ship it as a Spanish PHQ-9.
Where the item-level pattern matters most
A few patterns worth reading off the item responses directly, beyond the total:
- Somatic-dominant PHQ-9: high scores on sleep, appetite, energy, concentration; low scores on mood and self-worth. Often points at a medical workup that has not happened yet: thyroid, sleep apnea, anemia, B12.
- Mood-and-cognition-dominant PHQ-9: high scores on mood, self-worth, and anhedonia; lower somatics. More likely a primary depressive episode rather than secondary.
- Anhedonia plus item 9: even at moderate totals, this combination changes the conversation. Flat affect with passive suicidality is a different clinical situation than agitated depression with active SI.
- GAD-7 with high "trouble relaxing" and "afraid something awful might happen": anticipatory anxiety dominant, often more responsive to behavioral interventions than items dominated by restlessness or irritability.
None of these are diagnoses. They are reading-the-form patterns that change which questions you ask in the next ten minutes of the visit.
What the patient should see
Less than you think. The patient should see the questions, exactly as written, and a confirmation that their response was recorded. They should not see their score plotted on a chart against population norms. They should not see a severity band labeled "moderately severe depression" with no clinician interpretation.
The severity band, the trajectory, the item-level pattern: those belong in the clinician's review surface, not in the patient's hand. A patient who sees their PHQ-9 climb three points week-over-week without a clinician's framing is more likely to spiral than to act on it. The instrument was not designed to be self-interpreted, and we should stop pretending it was.
What we do in Nyra
PHQ-9 and GAD-7 are administered on days 1, 7, and 14 of the 14-day arc with the original wording, original order, original response options. Item 9 escalates the screen immediately with 988 and 911 as plain text. Item-level responses, severity band, and trajectory all show up in the clinician's evidence map alongside the patient's reflections, never in the patient's view. The patient sees that their response was saved and that their clinician will see it next time.
That is the whole product, on the scale side. The instrument does the work; we just deliver it without breaking it.
Where to go next
If you administer PHQ-9 / GAD-7 in your clinic and want to see how a 14-day pre-visit arc would land in your workflow, book a thirty-minute walkthrough. The demo includes the clinician evidence-map surface and the audit log we use to keep instrument integrity intact.
For deeper reading: our post on item-9 escalation design covers the safety surface in more detail, and the clinical instruments documentation lays out the full validated-wording policy.