通过统一的 trace_id 将浏览器错误摘要与后端 API 失败记录关联,是追踪 AI Agent 成本去向的有效方案。边界在于:后端错误 pipeline 作为系统记录源,前端仅转发紧凑摘要,完整源码映射和分布式追踪仍需专用工具。
The important trade-off is fidelity versus operational weight: use a backend error pipeline as the system of record for API failures and AI-call cost, then forward compact browser error summaries through that backend with the same trace_id or request_id. Short answer: this is the simplest defensible setup for a logistics agent loop when the question is "which failed shipment-planning run consumed this model cost?" It is not a substitute for source-map decoding, session replay, or a distributed trace viewer; add a browser-focused product when those are part of the debugging objective.
That boundary matters. A browser exception saying "plan generation failed" is nearly useless unless an operator can connect it to the API request, the model invocation, the carrier lookup, and the cost metadata recorded for that run. The correlation identifier provides the join key. It does not magically provide a span tree.
The evidence packet for one dispatch attempt
Consider a bounded failure during a dispatch window: the agent proposes no route, the browser reports an error, the API returns a failure, and several model calls have already accrued cost. My first move in the review would be to resist counting four alerts as four incidents. I would ask for one correlation value carried from the edge request into every backend log and into the sanitized browser report, then group the evidence by agent run.
The invariant is plain: one user-visible attempt needs one stable join key, while every internal operation may have its own span identifier. A trace_id follows the full attempt; a span_id distinguishes work such as model inference or a carrier API call; a request_id can remain the narrower HTTP identifier if that is already the platform convention. Pick the semantics once. Mixing those names for the same value produces correlations that look convincing and are wrong.
For an AI loop, record cost at the call boundary rather than estimating it later from an error count. Useful event fields include the stable trace value, agent run ID, model and vendor, cost_usd, latency_ms, outcome, and a low-cardinality stage such as classify, plan, or validate. Do not put prompts, access tokens, customer addresses, or arbitrary exception text into metric labels. Logs can carry carefully redacted detail; metrics should support capacity planning without creating an unbounded series count.
One short incident can otherwise inflate three numbers independently: browser errors, API exceptions, and failed model calls. The join key lets support reconstruct the sequence, while the agent run ID lets finance attribute spend to the workflow. Keep both.
How should a React frontend plus Node.js backend correlate errors?
The backend should issue or accept a valid correlation value, return it in a response header, and include it in structured logs. The browser's global error handler can read the value retained from the relevant API response and POST a sanitized summary to an application-owned endpoint. That endpoint, not the browser, sends data onward to the chosen error service, so no ingestion credential is exposed in client code.
This runnable Go service demonstrates the server-side contract. It uses only the standard library, applies a size limit, rejects malformed input, and logs JSON that can be joined on trace_id. It also exposes an internal handler that performs a real authenticated log search against a plain REST API, with an explicit method, status checks, and bounded rate-limit retries. No search filters appear because that route's discovery parameters are undeclared; inventing a convenient trace_id query parameter would make the sample look better while teaching an unsupported contract. The browser capture code is intentionally omitted because this article's code convention is Go-only; the wire contract is the part that must remain stable across frontend frameworks.
package main
import (
"crypto/rand"
"encoding/hex"
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"os"
"strconv"
"strings"
"time"
)
type browserError struct {
TraceID string `json:"trace_id"`
Message string `json:"message"`
Page string `json:"page"`
}
func newTraceID() (string, error) {
b := make([]byte, 16)
if _, err := rand.Read(b); err != nil {
return "", err
}
return hex.EncodeToString(b), nil
}
func agent(w http.ResponseWriter, r *http.Request) {
traceID := r.Header.Get("Traceparent")
if traceID == "" {
var err error
traceID, err = newTraceID()
if err != nil {
http.Error(w, "trace allocation failed", http.StatusInternalServerError)
return
}
}
w.Header().Set("X-Trace-ID", traceID)
log.Printf(`{"level":"info","event":"agent_request","trace_id":%q}`, traceID)
w.Header().Set("Content-Type", "application/json")
fmt.Fprintf(w, `{"status":"accepted","trace_id":%q}`, traceID)
}
func captureBrowserError(w http.ResponseWriter, r *http.Request) {
defer r.Body.Close()
var event browserError
decoder := json.NewDecoder(http.MaxBytesReader(w, r.Body, 16<<10))
decoder.DisallowUnknownFields()
if err := decoder.Decode(&event); err != nil || event.TraceID == "" || event.Message == "" {
http.Error(w, "invalid error report", http.StatusBadRequest)
return
}
log.Printf(`{"level":"error","event":"browser_error","trace_id":%q,"message":%q,"page":%q}`,
event.TraceID, event.Message, event.Page)
w.WriteHeader(http.StatusAccepted)
}
func searchLogs(ctxRequest *http.Request, baseURL, apiKey string) ([]byte, error) {
client := &http.Client{Timeout: 10 * time.Second}
url := strings.TrimRight(baseURL, "/") + "/v1/logs/search"
for attempt := 0; attempt < 4; attempt++ {
req, err := http.NewRequestWithContext(ctxRequest.Context(), http.MethodGet, url, nil)
if err != nil {
return nil, err
}
req.Header.Set("Authorization", "Bearer "+apiKey)
resp, err := client.Do(req)
if err != nil {
return nil, err
}
body, readErr := io.ReadAll(io.LimitReader(resp.Body, 1<<20))
resp.Body.Close()
if readErr != nil {
return nil, readErr
}
if resp.StatusCode == http.StatusTooManyRequests {
delay := time.Duration(1<<attempt) * time.Second
if seconds, err := strconv.Atoi(resp.Header.Get("Retry-After")); err == nil && seconds >= 0 {
delay = time.Duration(seconds) * time.Second
}
select {
case <-time.After(delay):
continue
case <-ctxRequest.Context().Done():
return nil, ctxRequest.Context().Err()
}
}
if resp.StatusCode < 200 || resp.StatusCode >= 300 {
return nil, fmt.Errorf("log search failed: status=%d body=%s", resp.StatusCode, body)
}
return body, nil
}
return nil, fmt.Errorf("log search remained rate limited")
}
func internalLogSearch(baseURL, apiKey string) http.HandlerFunc {
return func(w http.ResponseWriter, r *http.Request) {
body, err := searchLogs(r, baseURL, apiKey)
if err != nil {
http.Error(w, err.Error(), http.StatusBadGateway)
return
}
w.Header().Set("Content-Type", "application/json")
w.Write(body)
}
}
func main() {
baseURL := os.Getenv("INFRAI_BASE_URL")
apiKey := os.Getenv("INFRAI_API_KEY")
if baseURL == "" || apiKey == "" {
log.Fatal("INFRAI_BASE_URL and INFRAI_API_KEY are required")
}
mux := http.NewServeMux()
mux.HandleFunc("POST /agent/run", agent)
mux.HandleFunc("POST /client-errors", captureBrowserError)
mux.HandleFunc("GET /internal/log-search", internalLogSearch(baseURL, apiKey))
server := &http.
In production, validate an incoming W3C traceparent header instead of accepting arbitrary text, or generate the value at a trusted edge. Return a separate, simple X-Trace-ID if support staff need a copyable identifier. The browser report should contain the message, page, release, timestamp, and correlation value after redaction; stack collection is useful, but without source-map processing a minified stack will remain difficult to interpret.
There is also a retry trap. A browser may resend after a timeout even though the server accepted the first report, so give reports a client-generated event ID and deduplicate them at ingestion. This is less glamorous than a trace visualization, but it prevents one flaky connection from becoming five apparent failures.
The ownership test before choosing a product
The decision is not "which observability vendor is best?" It is which evidence must be available during the dispatch SLO's response window, and how much on-call machinery the platform team will own.
This table is deliberately unfair to anyone seeking a universal winner. There isn't one. Sentry is the sharper default when a minified browser stack must resolve to source and replay is operationally important. Datadog fits an organization already standardizing infrastructure, APM, logs, and real-user monitoring under one control plane. Honeycomb is compelling when engineers need to ask unplanned questions of richly structured events. OpenTelemetry is the sensible instrumentation layer when portability is a roadmap requirement, although the collector is not an incident-management product by itself.
The lighter REST option fits a narrower operating model: backend and API error capture, searchable logs, and manual correlation using stored identifiers. It can receive frontend summaries through your server. Because it has no threshold, phone, SMS, or webhook notification route, a team must poll the query surface and operate its own alert dispatcher; because it has no synthetic check or heartbeat monitor, use a service such as Healthchecks.io for silent "the dispatch job never ran" failures. Those are material on-call costs, not footnotes.
Eight events before the browser retries
Start with the error budget. If the dispatch agent has a service-level objective of 99.9% successful runs, one million monthly runs leave an error budget of 1,000 unsuccessful runs. That arithmetic is illustrative capacity planning, not a claim about any product's uptime. Define whether a run that returns a usable plan after one model retry is successful, degraded, or failed before dashboards begin making the decision for you.
Then size ingestion from events per run rather than from average requests per second. Suppose the application deliberately records one run summary, up to six model-call events, and one terminal error event. At 1,000,000 runs, the upper planning bound is 8,000,000 events per month before browser duplicates and retries. Peak dispatch periods matter more than the monthly average, so load-test the ingestion path at the burst rate and verify what happens under backpressure.
Cost attribution should follow the same hierarchy used for reliability: tenant or business unit, workflow, agent run, call. Record the provider-reported call cost where available, then roll it up; do not infer spend from latency, token guesses, or the presence of an exception. A failed loop may contain successful paid calls, and a successful loop may be wasteful because it retried five times.
This is where a superficially simple setup can become expensive to operate. Polling for errors needs a schedule, pagination discipline, a durable cursor, deduplication, and its own dead-man check. If the team cannot commit to testing that path and paging on its failure, buy alert delivery from a platform that owns it.
Two different kinds of silence
Use the backend-first pattern when support mainly needs to correlate a user-visible failure with API logs and AI-call cost, and when manual trace lookup is acceptable. It also works as a transitional architecture: preserve W3C trace context now, then add a trace backend later without changing the join semantics.
Do not use it alone for a consumer-facing browser application where minified stack decoding, release health, breadcrumbs, or session replay materially reduce mean time to recovery. Do not pretend identifier search is distributed tracing. A list of logs with matching text cannot show parent-child timing, missing spans, or the critical path.
Finally, separate active failure from silent absence. Error capture observes something that happened. A missed logistics reconciliation job emits nothing, which is why a heartbeat monitor belongs outside the job and outside the same failure domain. Three tools can be simpler than one improvised platform when each has a crisp responsibility.
The decision rule I would put in the roadmap is blunt: choose a dedicated browser monitor if frontend diagnosis drives the incident; choose a full observability suite or OpenTelemetry-backed stack if cross-service trace exploration drives it; choose the light REST pipeline if backend error capture, AI-call cost attribution, and manual ID correlation satisfy the SLO. Revisit the choice when manual joins consume more response time than the integration originally saved.
OpenTelemetry documentation
Sentry JavaScript source maps
Sentry Session Replay
Datadog Real User Monitoring
Honeycomb observability documentation
Healthchecks.io documentation
RFC 5424: The Syslog Protocol