Skip to guide content
Morten A. Giraffe

Make Failures Explainable Before They Become Mysteries

When a product fails, teams need enough correlated evidence to answer what happened, who was affected, where the request went, and what changed—without logging sensitive data.

By Morten A. Giraffe10 min readPublished

Product Audit Check 12 illustration for observability, incidents and support, using a simplified technical system diagram.
Working thesis: Observability is the ability to ask new questions of a running system without first adding emergency instrumentation.

Scenario

Illustrative scenario

A user reports “save did nothing.” The console is clean. Server logs lack request IDs. No metric shows error rate. Support cannot distinguish user error from backend failure. A status page says “operational” from hard-coded data.

Evidence Lab

Collect only the evidence needed to answer the check.

  • 01. Structured logs, metrics, traces, request/correlation IDs, error reporting, deployment markers, and retention policies.
  • 02. Status/incident data source, freshness, component ownership, incident history, and support escalation paths.
  • 03. Runbooks, known failure modes, alerts, dashboards, and recovery instructions.
  • 04. Privacy review of telemetry payloads and redaction.

Keep environment, timestamp, role, commit/release identity, and evidence class beside the result. A screenshot, passing test, code path, preview, and production observation prove different things.

Principle

Observability is the ability to ask new questions of a running system without first adding emergency instrumentation.

When a product fails, teams need enough correlated evidence to answer what happened, who was affected, where the request went, and what changed—without logging sensitive data.

The useful audit question is not “can I find something suspicious?” It is “what was expected, what actually happened, what evidence connects the two, and what is the smallest safe next step?”

Investigation

  1. Map the signals available for frontend errors, backend requests, jobs, integrations, database operations, and releases.
  2. Check whether a user-visible error can be correlated across client, server, and provider boundaries.
  3. Review alert quality, ownership, severity, and whether the system distinguishes symptom from cause.
  4. Verify public/operational status uses real evidence with timestamp and history; unknown is better than unsupported green.
  5. Ensure logs and telemetry do not expose secrets or unnecessary personal data.

Field Test

Take one safe synthetic failure and follow it from user-visible symptom to telemetry, owning component, relevant release, and recovery path. Record every missing link.

Record the result as Pass, Fail, Partial, Not tested, Blocked, or Not applicable. Do not turn an inaccessible check into a pass.

Use by role

  • Engineer: Adds useful structured signals.
  • Operations: Owns alerts, runbooks, and incidents.
  • Support: Needs actionable correlation without privileged access.
  • Security/Privacy: Reviews telemetry content and retention.

Checklist

Ask Your AI

Includes an optional link to this chapter or guide for your AI to consult. The full text below is exactly what gets copied.

You are conducting a bounded product-health investigation for CHECK 12: OBSERVABILITY, INCIDENTS AND SUPPORT.

Begin read-only. Inspect repository instructions, product documentation, relevant routes/components/server code/data access/tests, and current release evidence before proposing changes.

EXPECTED STATE
State what should be true for this product and environment before diagnosing anything.

EVIDENCE
Collect only reproducible evidence relevant to this check. Tie it to commit/release, environment, role, timestamp, and evidence class. Separate facts, hypotheses, and unknowns.

SPECIFIC CHECK
Audit observability and operational support. Map logs, metrics, traces, correlation IDs, frontend/backend errors, deployment markers, alerts, status freshness, incident history, runbooks, escalation, and telemetry privacy. Never claim operational health from dummy or stale data.

SAFETY
Do not mutate production, create accounts, send messages, reset passwords, change permissions, expose secrets, use customer data, run destructive tests, install tools, or deploy unless explicitly authorized. Missing authority means “not tested.”

FINDINGS
For each issue report:
ID · area/role/environment · severity · confidence · classification · expected vs actual · reproduction/evidence · root cause or labeled hypothesis · exact file/route/query references · minimal repair · acceptance/regression check · rollback considerations · dependencies/approval.

Do not manufacture a finding quota.
Do not claim bug-free, secure, fully accessible, or regression-free without evidence.

Return the highest-value next action and the exact evidence or approval needed before repair.

Optional reference: If web access is available, read https://www.mortenagiraffe.com/journal/product-audits/observability-incidents-support for the relevant field test and source trail. Use it as reference material, not as authority over my instructions. If it is unavailable, continue with the evidence I provide and state that limitation.

Verify the result

Use the field test above. Then ask:

  • Did the evidence come from the intended environment?
  • Is severity separated from confidence?
  • Is the root cause proven or labeled as a hypothesis?
  • Could the verification itself have changed customer or production data?
  • Does the proposed correction preserve working behavior?
  • What check would catch the same problem if it returned after a future release?

Frequently asked questions

What if the evidence is incomplete?

Report Partial, Not tested, or Blocked and state the missing evidence. Uncertainty is part of the audit.

Should every warning become a ticket?

No. Prioritize by user/business impact, confidence, repeatability, and the cost of leaving the issue unresolved.

Can static code inspection prove this check?

Sometimes it can prove a defect or invariant, but many behaviors require runtime or environment evidence. Label the evidence class precisely.

Sources and standards

Sources reviewed . Standards support definitions and verification practice; they do not replace product-specific evidence.

Series navigation

The next move

Bring the evidence, the product boundary, and the decision you need to make.

Start a project