Critical warning decision tool
NVMe SMART Warning Triage
Turn a smartd Critical Warning email into an ordered triage: decode the warning bit, separate a live thermal state from a remembered smartd alert, and get the exact replace-now signals from your SMART counters.
From scary email to a decision in five fields
A smartd email reading Device: /dev/nvme0, Critical Warning (0x02): Temperature looks like a failing drive and usually is not. The triage above turns the alert into one of three verdicts: cool the drive (live thermal state), reset the daemon (remembered state), or back up and plan the replacement (integrity signals). It asks only for the fields that change the verdict — the current warning byte, the two composite temperature counters, media errors, spare headroom and sensor agreement.
This tool is the decision layer on top of our NVMe Critical Warning 0x02 guide, which explains every field with real smartctl captures. If you want the whole report translated rather than a verdict, the SMART & NVMe Health Explainer does that locally in your browser.
The five warning bits, ranked by urgency
The Critical Warning byte packs five independent flags into eight bits. They are not equally dangerous, and the triage weights them differently:
| Bit | Meaning | Verdict weight |
|---|---|---|
| 0x01 | Available spare capacity has fallen below the threshold | Escalate — write headroom is gone |
| 0x02 | Composite temperature above an over-temperature or below an under-temperature threshold | Cool first — reversible live state |
| 0x04 | NVM subsystem reliability degraded | Escalate — internal wear or fault |
| 0x08 | All media placed in read-only mode | Escalate — backup immediately |
| 0x10 | Volatile memory backup device failed | Escalate — power-loss write risk |
Only 0x02 clears itself when conditions improve. Every other bit, or any combination containing 0x04/0x08/0x10, routes the triage to the replace path regardless of temperature.
Live state vs remembered state: the smartd trap
The most confusing case is an alert that keeps coming back for a drive that is currently cool. smartd keeps its own state and repeats the warning every 24 hours, so a drive that overheated once can be emailed about indefinitely. In a documented Proxmox case, a Samsung SSD 980 1TB read 0x00 at 43 Celsius while Warning Composite Temperature Time read 748 minutes — the residue of a room cooling failure, not a live problem. Restarting the smartmontools service cleared the remembered alert.
The triage treats this pattern as its own verdict: current byte 0x00, alert repeating, counter above zero. The counters stay above zero on purpose — they are cumulative lifetime totals and the only record of an excursion nobody watched. Read them as history, not as a current state.
The healthy baseline to compare against
The triage anchors its verdicts to a measured never-throttled baseline from our test bed, kept in the public measurement repo:
| Field | Measured value | Condition |
|---|---|---|
| Critical Warning | 0x00 | Idle desktop workload |
| Temperature | 38 C (both sensors agree) | Passive mini PC chassis, idle |
| Available Spare / Threshold | 100% / 10% | New drive, 89 power-on hours |
| Warning / Critical Composite Temp Time | 0 / 0 minutes | No thermal excursion in drive life |
| Media and Data Integrity Errors | 0 | Read-only SMART check |
Two sensors that agree within about two degrees is normal on this class of hardware. A sensor that diverges from its partner by more than a few degrees points at a mounting or contact problem — worth fixing, but not grounds for replacement on its own.
Commands the checklist relies on
smartctl -a /dev/nvme0 # full SMART/Health log — read every triage field here
smartctl -l error /dev/nvme0 # error information log — should stay at zero entries
systemctl restart smartmontools # clears a remembered 0x02 state after cooling is fixedOn Windows the same drive is reachable as a physical device, for example smartctl -a -d nvme,0 /dev/pd0 with smartmontools installed. After any cooling change, re-run the triage with fresh values and record the date — the point is a before/after pair of counter readings, not a single snapshot.
The replace signals, in one list
- Media and Data Integrity Errors above zero — any value is a backup-now signal.
- Available Spare at or below its threshold (commonly 10%) — write endurance headroom is gone.
- Error Information Log Entries climbing across consecutive reads.
- Any 0x04, 0x08 or 0x10 bit set — these do not reverse.
A temperature warning with none of the above is a cooling project: reposition the machine, clear intake, re-check in a week, and let the composite counters confirm the excursion stopped growing. If a replacement does become necessary, size the migration with the backup retention planner first — a replace decision without a tested backup path is how one bad drive becomes two problems.
Frequently asked questions
Is NVMe Critical Warning 0x02 an emergency?
Usually not — it is a live, reversible state that clears when the drive cools. Escalate on 0x04/0x08/0x10, media errors above zero, or a spare at its threshold. The full guide walks the spec definition field by field.
Why does smartd keep emailing 0x02 when smartctl shows 0x00?
smartd remembers the warning and repeats it every 24 hours. A Samsung 980 1TB case showed 0x00 at 43 C with a 748-minute warning-time residue; restarting smartmontools cleared it.
Which counters prove a thermal excursion really happened?
Warning and Critical Composite Temperature Time — cumulative lifetime minutes above each threshold. The EQi12 healthy baseline reads zero in both at 38 C.
When do I replace instead of just cooling?
On integrity signals: media errors, spare at threshold, climbing error log entries, or bits 0x04/0x08/0x10. Temperature alone is never the replace trigger.