Incident investigation best practices: from timeline to CAPA
A step-by-step methodology — secure the scene, build a timeline, find root causes, write corrective actions that stick.
QEHS Ethos Team
Founding team
The QEHS Ethos Team built the QEHS platform after a decade managing EHS programs in heavy industry. We write about safety culture, regulatory strategy, and how software can get out of the way.
12 min read
Reviewed by QEHS Ethos Team — Founding team
A poor investigation blames the operator and closes with "retrain all." A good investigation finds system-level root causes and produces SMART corrective actions that change the system. Steps: secure the scene, build a timeline, identify root causes (Five Whys, Fishbone), write SMART corrective actions, and verify effectiveness.
A QEHS platform supports investigation quality with structured RCA templates, auto-populated timelines, CAPA linking with SLAs, and immutable evidence chains. See the incident reporting use case, the near-miss guide, and the CAPA glossary entry.
The timeline before the causes is the principle that separates a real investigation from a rationalization, and it is the one the platform enforces. The first act of an investigation is to build the sequence of events from the evidence — the timestamps, the logs, the witness statements — without assigning cause, and the sequence that is built from the evidence is the sequence that holds up. The investigation that names a cause before it builds the timeline is the investigation that fits the evidence to the cause, and the cause that is named first is the cause that is usually the operator. The platform that auto-populates the timeline from the records is the platform that makes the evidence-first order the easy order, and the timeline that is built from the records is the timeline the investigator cannot revise to fit a theory.
The root cause is the thing the Five Whys and the Fishbone are built to find, and it is the thing the retrain-the-operator answer almost never is. The Five Whys ask why, repeatedly, until the answer stops being a person and starts being a system — the guard was bypassed because the access was inconvenient because the layout was changed without a review — and the Fishbone organizes the causes by category so the investigator does not fixate on the first one. The investigation that stops at the operator is the investigation that produces a corrective action the operator cannot carry out, and the corrective action that is retrain-all is the corrective action that does not change the system that produced the incident.
The SMART corrective action is the one that survives the closure, and it is the one the platform tracks to a verified end. A corrective action that is Specific, Measurable, Assignable, Realistic, and Time-bound is an action a person can complete and a system can verify; an action that is improve training is an action no one can close and no one can verify. The platform that links the corrective action to the investigation, gives it an owner and a due date, and requires the effectiveness check before the closure is the platform whose CAPA count is a count of finished work and not a count of good intentions. The closure without the effectiveness check is the closure that lets the same incident happen again, and the audit that reads the verified-CAPA rate is the audit that reads the quality of the investigation.
- Secure the scene and preserve the evidence — photographs, the equipment state, the logs — before the timeline, because the evidence that is lost in the first hour is the evidence the investigation cannot recover.
- Build the timeline from the records and the witness statements, without assigning cause, so the sequence is built from the evidence and not from a theory.
- Identify the root causes with the Five Whys or the Fishbone, and stop when the cause is a system and not a person, so the corrective action changes the system.
- Write the corrective actions in the SMART form, with an owner and a due date, and link each to the root cause it addresses, so the closure is a finished action and not an intention.
- Verify the effectiveness of each closed action against the evidence — the guard is back, the layout is reviewed, the near-miss rate dropped — before the investigation is closed, so the closure is a verified end.
The near-miss-to-investigation ratio is the leading indicator that tells the safety team whether the investigations are keeping up with the learning, and it is the one the platform reports. A site that investigates every near-miss is a site that learns before the incident, and a site that investigates only the recordable is a site that learns after the harm. The platform that holds the near-miss and the investigation on one tenant is the platform where the ratio is a live number and not a quarterly reconstruction, and the ratio that is rising is the one that shows the investigations are not keeping up with the reports. For the related practice, see the incident reporting use case, the near-miss guide, and the CAPA glossary entry; for the handoff from the near-miss to the corrective action, the near-miss to CAPA handoff post.
The investigation team is the composition that decides whether the investigation finds the system cause or stops at the operator, and it is the one the platform convenes. A team that has only the supervisor and the operator is a team that produces the retrain-all answer; a team that adds the engineer, the maintainer, and the safety lead is a team that finds the guard-bypass cause. The platform that pulls the team from the roles on the tenant is the platform that convenes the right team by default, and the team that is convened by the platform is the team that does not depend on the supervisor remembering who to call.
The evidence chain is the integrity the investigation depends on, and it is the one the platform guarantees by construction. A photograph, a log, a statement that is attached to the investigation record at the time it is captured is evidence the investigation can rely on; the same artifact collected in a separate system and attached later is evidence a defense can challenge. The platform that writes every artifact to the investigation record with the user and the timestamp, and that makes the record immutable after the closure, is the platform whose evidence chain holds up in a hearing, and the evidence chain that holds up is the one that turns an investigation into a defensible record.
The severity and the depth of the investigation are the two a program has to match, and the one a platform enforces with a grade. A near-miss that could have been a fatality deserves the depth of a fatality investigation, and a first-aid case does not; the program that investigates everything at the depth of the first-aid case is the program that misses the serious one, and the program that investigates everything at the depth of the fatality is the program that burns out the investigators. The severity grade that sets the investigation depth is the control that matches the effort to the risk, and the platform that routes by severity is the platform that does not ask the supervisor to decide the depth on the day of the event.
The investigation cadence is the one that keeps the learning from stalling, and it is the one the trend review surfaces. A site that investigates the incidents as they occur and never reviews the trends is a site that learns one at a time and never in the aggregate, and the pattern that is visible across ten investigations is the pattern that is invisible in any one. The platform that holds the investigations on one tenant is the platform where the trend review is a query, and the root-cause category that dominates the trend is the one the program addresses once instead of ten times. The cadence that investigates promptly and reviews periodically is the cadence that turns the investigations into a program and not a series.
The recordability intersection is the one the investigation shares with the recordkeeper, and it is the one the platform holds on one record. An incident that is investigated for its cause is also a case that may be recordable on the 300 log, and the two records on two systems are two records that describe one event. The platform that holds the investigation and the 300-log entry on one record is the platform where the recordability decision and the root-cause analysis are made against the same facts, and the two records that reconcile are the two records that hold up under an audit and a hearing.