◀ Knowledge hub

08, Observability, docs & ops

Production failures that arrive with a stack trace

codeAmani Labs Engineering
Cinematic still for Production failures that arrive with a stack trace

The errors you do not see are the ones that cost you

A production error that nobody is told about is a silent failure: the user hits it, gives up, and you find out weeks later from churn you cannot explain. Error tracking turns that silence into a signal. The moment something throws in production, you get the stack trace, the context, and a count of how many users it touched, before the support tickets start.

Capture the trace, not just the symptom

A log line that says "something went wrong" is almost useless. What you need is the exception with its full stack, the request that triggered it, and the state around it. Sentry captures that automatically when you initialize it, so an error arrives as a thing you can actually debug rather than a vague report.

import * as Sentry from "@sentry/nextjs";

Sentry.init({
  dsn: process.env.SENTRY_DSN,
  tracesSampleRate: 0.1, // sample performance, capture all errors
});

Add the context that makes triage fast

An error is much cheaper to fix when you know who hit it and what they were doing. Attach the user's id, the tenant, and the relevant identifiers, taking care never to attach sensitive data such as protected health information or secrets. The goal is enough context to reproduce, not a copy of the payload.

Group, prioritize, and act on volume

Sentry groups identical errors so a thousand occurrences of one bug show up as one issue with a count, not a thousand alerts. That count is your priority signal: a new error hitting many users jumps the queue over an old one that fires once a week. Wire alerts to the cases that matter and resist alerting on everything, because an alert channel that cries wolf gets muted, and a muted channel is the same as no channel.

Close the loop

An error tracker is only useful if errors get fixed and verified fixed. Tie the issue to the change that resolves it, and let the tracker confirm the error stopped occurring after the deploy. This is the observability layer doing its real job: failures surface immediately, and the knowledge of what broke and how it was fixed compounds rather than evaporating into a chat thread.

Qualified conversation

Have a build to de-risk? Let's talk.

Tell us what you are building. We respond within two business days.