The author describes building Kept, a self-hostable post-purchase support agent for e-commerce, with reliability as a central goal. Because a non-deterministic language model sits at the product's core, they designed the complete tracing layer before implementing the main agent loop. The central architectural decision was to keep an in-process trace as the source of truth for machines, especially evaluations, while using Langfuse as a projection and human-facing sink. This avoids extra API calls, latency, and concurrency issues in evaluations, and keeps transport or presentation bugs separate from core agent behavior. The trace model has three span types: a turn containing model-call and tool-execution spans; causality between related operations is carried by a callId. Separate fields record what the customer saw, what the model saw, and what actually happened during tool execution. The lifecycle distinguishes in-progress, completed, error, and undetermined states. An unfinished span at trace end becomes undetermined rather than disappearing or being mislabeled as an error, while tool-result status remains distinct from the enclosing span status. The implementation uses typed TypeScript classes, an in-process trace array, an OpenTelemetry GenAI attribute map, an OTLP mapper, and a generic exporter with a Langfuse transport. A fabricated order-status conversation exposed an attribute-mapping bug: Langfuse showed empty tool arguments even though the in-process trace was correct. Because evaluations read the in-process object, the presentation defect could not create a false evaluation failure. The author identifies three benefits: observability forces production debugging questions early, a greenfield codebase makes vendor-neutral best practices easier to adopt, and the span tree provides a first sketch of the agent state machine and tool contract. The costs are later redesign, throwaway traffic scripts, and the difficulty of applying the approach to constrained systems. For existing codebases, the author suggests designing traces for each new AI feature, after incidents, or by comparing an ideal trace with production data.
