Tracing
Trace a request from Railway's edge through a service and into the services it calls. Tracing is switched on per service and per environment: with the set-service-tracing MCP tool, railway trace enable, or the Tracing setup drawer on the project's Traces tab. Read the result with list-traces and get-trace, or on the same tab.
Tracing is a preview feature. If the project has no Traces tab, the account needs Tracing enabled in Priority Boarding first.
What Railway records without code changes
- Edge and proxy spans. For every sampled request to a traced service's public domain, Railway's edge records a server span (method, path, status, cache result, upstream) and the regional proxy adds a span for its hop. The edge forwards a W3C
traceparentheader to the service and stampsx-railway-trace-idon the response. - Automatic instrumentation (OBI). A per-service switch that attaches eBPF probes to the service's Node.js, Go, Python, Ruby, or Java processes on the host. It exports server spans for incoming HTTP/gRPC requests, client spans for plaintext outgoing calls, and spans for database and cache protocols it decodes. No SDK, no redeploy.
Spans from inside the service only appear once the service exports them, through automatic instrumentation or an OpenTelemetry SDK. A service without a public domain never gets edge spans; only what it exports itself shows up, joined to traces other services propagate to it over the private network.
Enable tracing
Tracing is two switches on each service instance, so they are set per service and per environment. tracingEnabled makes the edge trace requests to the service and gives its next deploy the OpenTelemetry variables below; autoInstrumentationEnabled attaches eBPF probes to the service's processes and only does anything while tracing is on. There is no project-wide default and no sample rate: every client-facing request to a traced service is traced. Enabling a service in production says nothing about staging, so check and set each environment the user cares about.
Dashboard: open the Traces tab → Tracing setup. The drawer lists the services of the environment the dashboard is on (the subtitle names it); each row's Traced switch and Automatic instrumentation / Manual instrumentation choice apply to that environment only. Switch environment to configure another. The service's Settings → Tracing section links to the same drawer.
Agent path: two MCP tools read and change these settings, and railway trace does the same from the CLI. Resolve IDs from the URL or railway status --json first, and read before writing. Both tools use the production environment when environmentId is omitted, so pass it whenever the user is looking at another environment. On a Railway cloud agent in a dashboard chat session the railway CLI is unauthenticated, so resolve IDs with list-services and stay on the MCP tools throughout.
| Tool | Access | Purpose |
|---|---|---|
get-tracing |
viewer | The environment and, per service in it, tracingEnabled, autoInstrumentationEnabled and whether it is autoInstrumentationActive (both on). Pass serviceId for one service, omit it for every service in the environment |
set-service-tracing |
member | tracingEnabled and autoInstrumentationEnabled for one service in one environment. Each is optional and independent; omit what should stay as it is. Returns the service's state in that environment |
Both take projectId and an optional environmentId. describe-service reports the same tracing state for one service in the environment it was asked about.
Get tracing for project 6adb5ae3-0e3a-4ead-b42c-1fd36f217ffb in environment <environment-id>Set service tracing for project 6adb5ae3-0e3a-4ead-b42c-1fd36f217ffb, service <service-id>, environment <environment-id>: tracingEnabled trueSet service tracing for project 6adb5ae3-0e3a-4ead-b42c-1fd36f217ffb, service <service-id>, environment <environment-id>: autoInstrumentationEnabled trueset-service-tracing warns when it switches auto-instrumentation on for a service whose tracing is off: the switch does nothing until tracing is enabled too. Without Railway MCP or railway trace, the public serviceInstanceUpdate mutation sets the same fields through railway api, and the environment's service instances carry them for reading; see request.md. That fallback needs an authenticated CLI, which a cloud agent's chat session does not have.
railway api \
'mutation traceInstance($serviceId: String!, $environmentId: String!) {
serviceInstanceUpdate(serviceId: $serviceId, environmentId: $environmentId,
input: { tracingEnabled: true, autoInstrumentationEnabled: false })
}' \
--variables '{"serviceId":"<service-id>","environmentId":"<environment-id>"}'
railway api \
'query tracing($id: String!) {
environment(id: $id) { serviceInstances { edges { node {
serviceId serviceName tracingEnabled autoInstrumentationEnabled
} } } }
}' \
--variables '{"id":"<environment-id>"}'The older Project and Service tracing fields still exist as deprecated stubs: projectUpdate(tracingEnabled) switches every instance in the project, serviceUpdate(tracingEnabled) the service in every environment, and tracingSampleRate is ignored. Don't use them; they flip environments the user didn't mention.
CLI: railway trace (aliases traces, tracing) sets the same switches, in one environment at a time: the linked one, or --environment <name>. This needs a CLI release newer than 5.62.1. An older CLI still runs railway trace enable, but through the deprecated service field, so it sets the service in every environment, and its inherit, --project-default and --sample-rate belong to the model that is gone. Use the CLI when the user works in a linked repo or wants exact command output; on a cloud agent's chat session stay on MCP.
railway trace status # linked service in the linked environment
railway trace status --all # every service in the environment, last edge/app span
railway trace enable --service <service> # trace one service in the linked environment
railway trace enable --auto-instrument # tracing plus OBI for the linked service
railway trace enable --all --environment staging # every service in staging
railway trace disable --auto-instrument # tracing and OBI off--all and --service are mutually exclusive. Pass --project <id> --environment <name> together when nothing is linked. --json prints one document with environment: { id, name } and the service state. Changing tracing needs a user or workspace token; a project token (RAILWAY_TOKEN) can only read. The CLI has no switch for auto-instrumentation alone: railway trace disable --auto-instrument followed by railway trace enable leaves tracing on and OBI off, or use set-service-tracing with only autoInstrumentationEnabled.
What happens next:
- The edge starts tracing requests to the service's domains in that environment within seconds.
- Automatic instrumentation reaches the running containers within about a minute. No redeploy.
- The OpenTelemetry variables below are added on the next deploy in that environment. An app with an SDK exports nothing until it is redeployed:
railway redeploy --service <service> --yes.
Infrastructure as code
The environment config carries the same two switches as services[<id>].tracing: { enabled, autoInstrumentation }, so tracing is part of what railway config pull, plan and apply manage (see iac.md). Railway serialises only the switches that are on: an untraced service has no tracing block, and enabled: false is the same as leaving it out. railway config pull renders the block for a traced service:
const api = service("api", {
source: github("owner/repo", { branch: "main" }),
tracing: { enabled: true, autoInstrumentation: true },
});A plan shows a tracing change as a resource.update with field: "tracing". Turning enabled on or off has deployEffect: deploy, because the OTEL_* variables land with a deploy; an autoInstrumentation-only change is deployEffect: none and reaches the running containers live. A new service carries tracing on its create.
Current limitation: the SDKs don't carry tracing yet. The TypeScript service() helper, and the Python and Go ones, whitelist their config keys and tracing is not among them, so a tracing: the user writes in railway.ts is dropped before the CLI sees it and never applies. The other direction is worse: a compiled config with no tracing block asks to remove it, so railway config plan against a service that is traced shows a tracing update with after: null and deployEffect: deploy, and railway config apply turns tracing off and redeploys. Until the SDKs ship the field:
- Set tracing with
set-service-tracing,railway traceor the dashboard, not inrailway.ts, and say why when the user asks for it in config. - Read every plan for a
tracingchange nobody authored, and don't apply one.railway config pullreproduces the block, but the SDK drops it again on the next plan, so pulling doesn't fix it. - If an apply already disabled tracing, re-enable it with
set-service-tracing; the redeploy for the variables follows.
Choose how the service exports spans
Recommend the OpenTelemetry SDK. Spans the service exports itself are much higher quality than what automatic instrumentation can recover from outside the process: they carry the route, the query, the business identifiers and the errors the code knows about; they nest under handler-level spans, so the trace shows which step of a request took the time instead of a flat list of calls; they continue the trace through queue consumers, cron work and TLS callees; and they work in any language. Automatic instrumentation is the fallback for a service whose code can't be changed right now, or a first look while the SDK is being added. Present it as a stopgap, not the destination. When a user asks to "instrument my app", default to the SDK unless they ask for the no-code path.
| Automatic instrumentation (OBI) | OpenTelemetry SDK | |
|---|---|---|
| Code changes | None | Install the SDK, load it before the app serves requests |
| Takes effect | About a minute after enabling, no redeploy | Next deploy |
| Runtimes | Node.js, Go, Python, Ruby, Java. Not Bun, so not Functions | Any language with an OTel SDK |
| Captures | Incoming HTTP/gRPC, plaintext outgoing HTTP/gRPC, decoded DB and cache protocols | Whatever the SDK's instrumentations cover, plus custom spans and attributes |
| Misses | Outbound TLS callees don't join the trace; queue consumers, cron work and background jobs start new traces; Node.js and Python context propagation is best effort | Nothing structural; depends on the instrumentations you enable |
Pick one per service. Running an SDK in a service that also has automatic instrumentation produces duplicate spans for every request. When moving from OBI to an SDK, deploy the SDK first, confirm its spans arrive, then switch the service to manual instrumentation.
Instrument with an OpenTelemetry SDK
When tracing is on for a service in an environment, its next deploy there gets these variables. They show up in the service's Variables tab alongside the other Railway-provided variables and every OpenTelemetry SDK reads them, so an SDK configured without an explicit endpoint exports to Railway:
| Variable | Value |
|---|---|
OTEL_EXPORTER_OTLP_ENDPOINT |
Railway's OTLP receiver on the host running the service |
OTEL_EXPORTER_OTLP_PROTOCOL |
http/protobuf |
OTEL_EXPORTER_OTLP_HEADERS |
A header the receiver requires on every export |
OTEL_SERVICE_NAME |
The Railway service name |
OTEL_SERVICE_VERSION |
The commit SHA, or the deployment ID for image and CLI deploys |
Rules the agent must apply:
Don't hardcode the endpoint, protocol, or header in code or Dockerfiles. Let the SDK read the variables.
The receiver accepts traces only. Most SDKs also export metrics and logs to the same endpoint by default and log errors when that fails. Set both on the service with the
set-variablesMCP tool (skipDeploys: true, since the deploy that adds the SDK picks them up);railway variable set ... --skip-deploysdoes the same from a linked repo:Set variables for project <project-id>, service <service-id>: OTEL_METRICS_EXPORTER=none, OTEL_LOGS_EXPORTER=none, skipDeploys trueUser variables win. A service that sets its own
OTEL_EXPORTER_OTLP_ENDPOINTorOTEL_EXPORTER_OTLP_TRACES_ENDPOINT(for example to keep exporting to its own collector) gets none of the tracing variables, and its spans don't reach the Traces tab; edge spans still do. Railway sets no sampler variables, so a sampler the service configures is the only one in play. Checkrailway variable list --service <service> --jsonbefore assuming the provided values apply.Load the SDK first. Instrumentation libraries (
@opentelemetry/auto-instrumentations-node,opentelemetry-instrumentation-*and the like) patch libraries at import time, so the SDK must be loaded before the app's modules: a--require/--importflag in the start command,NODE_OPTIONS, a Pythonopentelemetry-instrumentwrapper, a Java-javaagent, and so on. Set it withrailway environment edit --service-config <service> deploy.startCommand "<command>"or as a variable; see deploy.md.Keep W3C Trace Context propagation on (the SDK default in most languages; Go requires setting the propagator explicitly) so the service continues the edge's trace instead of starting its own.
Don't override
OTEL_SERVICE_NAMEunless the user wants spans attributed under a different name than the Railway service.
Per-language install steps, framework notes, and a custom-span example are in the docs: Node.js, Deno, Functions (Bun), Python, Go, Java, Ruby, .NET, Rust, PHP. Fetch the page for the user's stack rather than reciting SDK commands from memory.
What to instrument
Instrumentation libraries give a trace its skeleton: a server span per request and a client span per call a library recognises. On its own that is barely better than automatic instrumentation. The value comes from spans the app opens itself around the units of work its authors think in. When adding tracing to a codebase, or when asked what to instrument, cover these three layers, in this order:
Inbound work. One span per HTTP handler, gRPC method, queue or job consumer invocation, cron run and WebSocket message. The HTTP and gRPC server instrumentations usually cover handlers. Consumers, cron and background jobs are not covered by the edge or by most instrumentations, so wrap each message or run in its own span (kind
CONSUMERorINTERNAL) and, where the message carries atraceparent, continue that context so the trace links back to the producer. Without this, background work is invisible or shows up as orphaned client spans.I/O: database queries, cache calls, outgoing HTTP and gRPC calls, queue publishes. This is where latency and failures hide. Use the driver or client instrumentation where one exists. Where none does (a raw socket client, a vendor SDK the instrumentations don't know), wrap the call in a
CLIENTspan with the standarddb.*,server.addressorurl.fullattributes. Every remote call should be visible as its own span. List every database driver or ORM, cache, HTTP and queue client in the dependency manifest and check each one, because the bundles miss common ones:- Node.js:
@opentelemetry/auto-instrumentations-nodecoverspg,mysql2,ioredis,mongodb,mongoose,knexandundici/http, but not Prisma (add@prisma/instrumentationto the SDK's instrumentations) and not postgres.js, Drizzle on postgres.js,Bun.sqlor the Neon serverless driver (no instrumentation exists; wrap queries in aCLIENTspan withdb.system,db.namespace,db.operation.name). An app written as ES modules needs the@opentelemetry/instrumentation/hook.mjsloader, or imported drivers stay unpatched while thehttpserver span still appears. - Python:
opentelemetry-distroalone instruments nothing; runopentelemetry-bootstrap -a installso the instrumentations for the installed drivers (psycopg2,asyncpg,SQLAlchemy,redis,pymongo,requests,httpx) are installed too. - Go: nothing is automatic. Wrap
database/sqlwithotelsql, pgx withotelpgx, go-redis withredisotel, HTTP clients withotelhttp.NewTransport. - Rust:
sqlxemits tracing events, not spans; wrap queries in a span.reqwestneedsreqwest-tracing. - Deno:
OTEL_DENO=truecoversfetch,Deno.serveandnode:http; npm database drivers need their own instrumentation or a hand-written span. - Java, Ruby, .NET, PHP: the agent,
opentelemetry-instrumentation-all, the per-library packages (Npgsql.OpenTelemetry,OpenTelemetry.Instrumentation.SqlClient, ...) andopentelemetry-auto-pdocover JDBC, ActiveRecord, ADO.NET and PDO; confirm the package for the driver in use is present. A server span with no client spans under it means this layer was skipped. Before opening the PR, run the app once with the console exporter (OTEL_TRACES_EXPORTER=consoleor the runtime's equivalent) and hit a route that queries the database: aCLIENTspan carryingdb.system, even a failed one, proves the driver is patched. List each client and the instrumentation that covers it, or why it is not covered, in the PR description.
- Node.js:
Logical units of work. A function that groups several I/O calls into one meaningful step (
checkout,syncUser,renderInvoice), or one that is CPU-heavy on its own (parsing a large payload, image resizing, template rendering, serialising a big response). Give each anINTERNALspan carrying the identifiers a person debugging it would want (order.id,tenant, item counts). This turns twelve siblingSELECTspans under a handler into a tree that reads like the code and says which step took the time. Rule of thumb: if you would want log lines saying "starting X" and "finished X in N ms", X is a span.
Keep spans out of tight loops and trivial helpers. A span per item in a loop of thousands hits the per-replica export limit (see Sampling) and adds nothing a count attribute on the parent wouldn't; instrument the loop, not the iteration. Name spans with low-cardinality names (GET /users/{id}, db.query users, process order) and put the variable parts in attributes. Set span status to ERROR and record the exception when a unit fails, so @status:error finds it. Never put secrets, tokens or raw personal data in span names or attributes.
Instrument a Function (Bun)
A Railway Function is a service whose source image starts with ghcr.io/railwayapp/function- (function-bun:1.4.0 today) and whose code is one TypeScript file, base64-encoded into the start command. get-service-config shows the image; get-function-source-code returns the code. Automatic instrumentation does not cover Bun, and Bun 1.4 has no OpenTelemetry of its own, so a function exports spans only through the OpenTelemetry JavaScript SDK loaded inside that one file.
What differs from a repo service:
- One file, no start command. There is no
--require,--preloadorbunfig.toml. Put the SDK setup at the top of the file. It runs beforeBun.servetakes its first request, which is all that is needed, because nothing gets monkey-patched. - Dependencies come from imports. The runtime turns every bare import into a
package.jsonentry and runsbun installat every cold start, without a cache. Pin withpkg@versionspecifiers:hono@4,@hono/otel@1,@opentelemetry/api@1, and@opentelemetry/sdk-nodeto the exact0.xversion tested, since it has no stable major. Every package added lengthens the cold start. - Nothing is instrumented for free.
NodeSDKconfigures the exporter, the resource and W3C propagation from theOTEL_*variables, but no OpenTelemetry package instrumentsBun.serve, Bun'sfetch,Bun.sqlorBun.redis(the Nodehttpandundiciinstrumentations don't see them). Incoming requests need@hono/otel(Hono) or a hand-written wrapper (Bun.serve); outgoingfetchcalls needpropagation.injectfor the callee to join the trace. The three layers in What to instrument still apply; the function's handlers, I/O and logical units are what to wrap. - A code push is a deploy. The variables land on the next deploy, and
update-function-source-codeorrailway functions pushis one, so a single push adds the SDK and picks up the variables.
Recipe:
Turn tracing on for the function in the environment it runs in with
set-service-tracing(orrailway trace enable --service <function>) ifget-tracingshows it off there. LeaveautoInstrumentationEnabledoff; it does nothing for Bun.Set the exporter variables with
set-variables.NodeSDKexports metrics and logs over OTLP by default and the receiver rejects both. PassskipDeploys: true; the code push in step 5 is the deploy that picks them up. Checklist-variablesfirst: a function that sets its ownOTEL_EXPORTER_OTLP_ENDPOINTgets none of Railway's tracing variables.Set variables for project <project-id>, service <function-id>: OTEL_METRICS_EXPORTER=none, OTEL_LOGS_EXPORTER=none, skipDeploys trueRead the code with
get-function-source-code:codeis what is current,deployedCodewhat runs,stagedwhether a commit is pending. Edit that, never a version recalled from memory.Add the SDK block at the top and wrap the requests, leaving the rest of the file as it is. A Hono function ends up like this:
import { NodeSDK } from "@opentelemetry/sdk-node@0.222.0"; import { trace } from "@opentelemetry/api@1"; import { Hono } from "hono@4"; import { httpInstrumentationMiddleware } from "@hono/otel@1"; // Reads OTEL_EXPORTER_OTLP_* and OTEL_SERVICE_NAME from the variables // Railway provides. Nothing to configure. const sdk = new NodeSDK(); sdk.start(); process.on("SIGTERM", () => sdk.shutdown().finally(() => process.exit(0))); const tracer = trace.getTracer("greeter"); const app = new Hono(); // One SERVER span per request, continuing the edge's traceparent. app.use(httpInstrumentationMiddleware()); app.get("/hello/:name", async (c) => { const name = c.req.param("name"); const greeting = await tracer.startActiveSpan("build-greeting", async (span) => { try { span.setAttribute("greeting.name", name); return `Hello, ${name}`; } finally { span.end(); } }); return c.json({ greeting }); }); export default { port: Number(Bun.env.PORT ?? 3000), fetch: app.fetch };For a function that calls
Bun.serveitself, wrap itsfetchhandler: take the parent frompropagation.extract(context.active(), req.headers, { get: (h, k) => h.get(k) ?? undefined, keys: (h) => [...h.keys()] })and run the handler insidetracer.startActiveSpan(name, { kind: SpanKind.SERVER }, parent, ...). For an outgoing call, start aSpanKind.CLIENTspan andpropagation.inject(context.active(), headers, { set: (h, k, v) => h.set(k, v) })into aHeadersobject beforefetch. A cron or script function has no server: its spans are new roots, recorded on every run, and it mustawait sdk.shutdown()as its last statement, or the batch never leaves the process. The docs page has all three in full.Write it back with
update-function-source-code(the whole file; passstaged: trueto stage instead of deploying live) orrailway functions push --path <file>. This deploy also adds the variables.Verify with the steps below:
curl -sI https://<domain>/hello/x | grep -i x-railway-trace-id, thenget-traceon the ID. A span withcomponentserviceand scope@hono/otel(or the tracer name) means the function exports. If the deploy logs showbun installfailing, an import specifier is wrong; if they show OTLP export errors for metrics or logs, step 2 was skipped.
Docs: Functions.
Sampling
Every request is traced. The edge traces 100% of client-facing requests to a traced service and writes the decision into the
traceparentsampled flag. Railway sets no sampler variables, so the SDK's default parent-based sampler follows the edge and records every root span of its own (cron jobs, queue consumers, private-network calls without atraceparent). There is no project or service rate to set.A client
traceparentoverrides. A request that arrives with the sampled flag set is always traced; one with it cleared never is. An instrumented client can start a trace that continues into Railway, and a caller can keep a request out of tracing by sending a cleared flag. To trace one specific request while debugging:curl -H "traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01" https://<domain>/<path>Generate a fresh random 32-hex trace ID (the second field) for each request; reusing one merges requests into a single trace.
Sample in the SDK when the volume is too high. Each replica can export 1,000 spans per 10 seconds; exports over the limit are rejected and the SDK reports a partial success. Disable noisy instrumentations first. If that isn't enough, set
OTEL_TRACES_SAMPLER=traceidratioandOTEL_TRACES_SAMPLER_ARG=<fraction>on the service yourself; aparentbased_*sampler follows the edge's flag and would change nothing for edge requests. Traces the SDK drops still show the edge and proxy hops.
Read traces
Traces are read through Remote MCP (the default agent path), railway trace list and railway trace get on the CLI, or the dashboard. The traces, trace and tracingStatus queries are on the public GraphQL API as well, so railway api can fetch them where neither fits.
| Tool | Access | Purpose |
|---|---|---|
list-traces |
viewer | Traces of an environment, newest first, one row per request with at least one span matching filter. Optional serviceId, startDate/endDate (ISO 8601 with timezone; defaults to the last hour), limit (default 100, max 500) |
get-trace |
viewer | One trace as an indented span tree (each line shows the span kind and the attribute that says what it talked to: db.system, http.route, url.full or server.address), with every span's attributes, events and links in the structured result. Takes traceId (32 hex characters); maxSpans caps the result and the output says when it was hit |
get-tracing-coverage |
viewer | What one service's own instrumentation covers: its spans in the window (default the last hour) by span kind, by the remote system they name (db.system, messaging.system, rpc.system, or http for HTTP client spans) and by the instrumentation scope that emitted them (@opentelemetry/instrumentation-pg, ...). Takes serviceId; optional startDate/endDate. Edge and proxy spans are left out. The one call that answers "are the database queries instrumented?" |
Both take projectId and an optional environmentId; omit it and the production environment is used, so pass the ID explicitly when the user is looking at another environment. Traces belong to the environment they were exported from.
List traces for project 6adb5ae3-0e3a-4ead-b42c-1fd36f217ffb in environment <environment-id> with filter "@status:error"Get trace 4bf92f3577b34da6a3ce929d0e0e4736 for project 6adb5ae3-0e3a-4ead-b42c-1fd36f217ffbfilter uses the same syntax as logs: @status:error, @component:edge AND @duration:>1000, @service:api AND @kind:client, @http.route:/checkout AND @http.response.status_code:500, @name:SELECT*. Built-in fields are trace, span, name, serviceName, service, deployment, replica, component (edge, proxy, service), kind, status, duration (ms); any other @key matches a span or resource attribute, free text matches the span name, and - negates. Narrow the filter or the window before raising limit.
On the CLI, list scopes to the linked service unless --all, --since/--until take relative (30m, 2h, 1d) or ISO 8601 times, --limit is 1 to 500 (default 100), --errors adds @status:error, and get --max-spans goes up to 2000. Human output is a table and an indented span tree; --json prints one trace summary or one span per line, like railway logs --json.
railway trace list --since 30m --errors --json
railway trace list --all --filter '@http.route:/api/users @duration:>500'
railway trace get 4bf92f3577b34da6a3ce929d0e0e4736 --jsonWorkflow for "why is this request slow / failing":
list-traces(orrailway trace list) with a filter that isolates the symptom (@status:error,@duration:>1000,@http.route:<route>), optionallyserviceIdfor one service.get-trace(orrailway trace get) on a returnedtraceId. The tree runs from the edge span down through every service; thecomponenton each span says which hop exported it, and a span withERRORstatus carries the message.- Read the span attributes in the structured result for the detail (
http.route,http.response.status_code,db.statement, custom attributes).
Verify tracing works
Prove the edge traced a request. The header is present only on traced responses:
curl -sI https://<domain>/ | grep -i x-railway-trace-idFetch that trace with
get-traceorrailway trace get <trace-id>and the returned ID. Edge and proxy spans confirm tracing is on; a span withcomponentserviceconfirms the app is exporting. After enabling an SDK, that only happens once the redeploy that added the variables is live; after enabling automatic instrumentation, allow about a minute.Check what the app covers with
get-tracing-coveragefor the service, after sending a few requests that hit the database. Every database, cache and queue the service uses must appear under "Remote systems" and the driver instrumentations under "Instrumentation scopes".SERVERspans with no database client span mean a driver is still unpatched (see the I/O layer in What to instrument);list-traceswith@kind:client @db.system:*is the same check as a trace list.Or watch the dashboard. The Traces tab is at
https://railway.com/project/<project-id>/traces?environmentId=<environment-id>; its Trace ID field accepts a bare 32-hex ID or a wholetraceparentheader. In Tracing setup, each service row shows when the edge and the app last exported a span, and the App indicator turns green on the first span from the service itself.
Troubleshoot
- No traces at all: confirm
get-tracingorrailway trace statusreports tracing on for the service in the environment the user is looking at (the switches and the traces both belong to an environment), the service has a public domain, and the account has Tracing in Priority Boarding. With little traffic, send a request and check forx-railway-trace-idon the response, thenget-traceit. - Traced in
productionbut not elsewhere: tracing is set per environment. Enable it in the other environment withset-service-tracingand itsenvironmentId, orrailway trace enable --environment <name>. - Edge spans only, nothing from the app (
get-traceshows onlyedgeandproxycomponents): the variables land on the next deploy, so redeploy. Then check the service doesn't set its ownOTEL_EXPORTER_OTLP_ENDPOINT, the SDK loads before the app serves, and, for OBI, the process is a supported runtime handling HTTP or gRPC. - Server spans but no database or client spans (
get-tracing-coverageshowsSERVERandINTERNALonly, orget-traceshows a request span with noCLIENTchildren): the driver is not instrumented. In Node.js an ES-module app without the@opentelemetry/instrumentation/hook.mjsloader, Prisma without@prisma/instrumentation, or a client no instrumentation covers (postgres.js, Drizzle on it,Bun.sql); in Python a missingopentelemetry-bootstrap -a install; in Go an unwrappeddatabase/sql. See the I/O layer in What to instrument. - App spans appear as separate traces (
list-tracesshows service-rooted traces withhasEdgefalse next to edge-only ones): the SDK isn't readingtraceparent. Enable the W3C Trace Context propagator and make sure nothing in front of the handlers strips the header. - SDK logs metrics or logs export errors: set
OTEL_METRICS_EXPORTER=noneandOTEL_LOGS_EXPORTER=none. - Duplicate spans per request: the service runs an SDK with automatic instrumentation on. Switch it off with
set-service-tracing(autoInstrumentationEnabledfalse), or on the CLIrailway trace disable --auto-instrumentthenrailway trace enable. - Spans missing from a busy service: over 1,000 spans per replica per 10 seconds. Disable noisy instrumentations or sample in the SDK; see Sampling.
- A Function shows edge spans only: automatic instrumentation can't help (Bun); the SDK has to be in the file. Check
get-function-source-codefor theNodeSDKblock and the request wrapper, that the deploy logs showbun installsucceeding, and thatOTEL_METRICS_EXPORTER/OTEL_LOGS_EXPORTERarenone. See Instrument a Function (Bun). railway config planwants to removetracing: the IaC SDKs drop the field, so every plan against a traced service proposes turning it off. Don't apply it; see Infrastructure as code.
Validated against
- Docs: tracing.md, automatic-instrumentation.md, nodejs.md, tracing/functions.md, functions.md, variables/reference.md, cli/trace.md, infrastructure-as-code/reference.md
- Public GraphQL API (
railway api schema):ServiceInstance.tracingEnabled,ServiceInstance.autoInstrumentationEnabledandServiceInstanceUpdateInputas the per-environment path;Project.tracingEnabled,Project.tracingSampleRate,Service.tracingEnabled,Service.autoInstrumentationEnabledand their update inputs marked deprecated; thetraces,traceandtracingStatusqueries - Railway MCP (
https://mcp.railway.com) tool descriptions and schemas:get-tracingandset-service-tracingwithenvironmentId,describe-service,list-traces,get-trace,get-tracing-coverage,get-function-source-code,update-function-source-code - CLI source (railwayapp/cli, per-environment
railway traceand the IaCtracingblock in #1239):src/commands/trace.rs(subcommands,--all,--environment),src/iac/compiler.rs(services[id].tracingon service nodes),src/iac/change_set.rs(tracing diff,deployEffect, removal when the block is absent),src/commands/config/mod.rs(pull renderer) - Provided variables observed on a traced service's deploy: the five
OTEL_*variables in the table above and no sampler variables - Function runtime: the public
ghcr.io/railwayapp/function-bun:1.4.0image (Bun 1.4.0, bare imports turned into apackage.jsonand installed withbun installat every start) - Function examples run on Bun 1.4.0 with
@opentelemetry/sdk-node0.222.0 and@hono/otel1.1.2 against a stub OTLP receiver: server span continues the incomingtraceparent, client span propagates it, script flushes onsdk.shutdown()