Profiling the native wado binary
Host-side Rust profiling (the compiler, wado serve, wado run, …
including wasmtime/cranelift). For the guest wasm program, use
wado-performance instead.
Pick a build profile
Choose based on what you're optimising for:
| Profile | Cargo flag | Use when | Trade-off |
|---|---|---|---|
profiling |
cargo build --profile profiling --bin wado |
Improving benchmark scores — release-equivalent codegen with debug info kept. Inherits release (thin LTO, codegen-units=1) and adds debug = 2, strip = false. |
Slow build (LTO), but the CPU profile reflects what users actually run. |
dev |
cargo build --bin wado |
Improving developer-iteration time — making cargo run -- compile/test/... faster for compiler-hackers. Uses the in-workspace [profile.dev.package.wado-compiler] opt-level = 1, so the compiler itself isn't molasses while everything else stays unoptimised. |
Fast build, but the absolute numbers are larger than a release run; ratios between hot paths are still actionable. |
If unsure: pick profiling for "users complain it's slow", pick dev
for "rebuild → run → tweak feels slow during development." The
analyzer script and recording flow below are identical for both.
Workflow
# 1. Build with the chosen profile (see table above)
cargo build --profile profiling --bin wado # benchmark-oriented
# or
cargo build --bin wado # dev-iteration-oriented
# 2. Record under load with samply (cargo install samply)
samply record --save-only --rate 1000 -o /tmp/prof.json -- \
target/profiling/wado serve --addr 127.0.0.1:8080 app.wado &
SAMPLY_PID=$!
# ... drive load (e.g. oha against benchmark/http_routing) ...
# 3. Stop: SIGTERM the CHILD, not samply. samply finalizes on child exit;
# signalling samply leaves the child running and the recording hangs.
kill -TERM "$(pgrep -P "$SAMPLY_PID" | head -1)"; wait "$SAMPLY_PID"
# 4. Analyze
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts /tmp/prof.jsonSymbols resolve against the path the profile recorded, not the binary that
produced it. Rebuilding over target/profiling/wado makes every earlier profile
symbolicate against the new build, and its output still looks normal. When you
A/B, give each arm its own path: copy the first build aside and profile it
there.
For one-shot commands (wado compile foo.wado, wado test foo.wado)
there is nothing to drive — samply records until the child exits, so
just invoke it directly:
samply record --save-only --rate 1000 -o /tmp/prof.json -- \
target/debug/wado test package-gale/tests/driver_rust_test.wadoInteractive call tree (and correct kernel symbols):
samply load /tmp/prof.json opens a browser-based call-tree UI. The
CLI analyzer below is for grep-able, transcript-friendly summaries.
The analyzer is a TypeScript script run directly by Node.js (>= 23.6,
which strips types with no flags). No build step or dependencies are
needed; node analyze_native_profile.ts ... just works.
Linux setup
# samply needs perf_event_paranoid <= 1 for a non-root user
echo '1' | sudo tee /proc/sys/kernel/perf_event_paranoid
# `addr2line` is part of binutils — usually already installed
addr2line --version >/dev/null || sudo apt-get install -y binutilsHow analyze_native_profile.ts works
samply's --save-only profile is unsymbolicated: funcTable.name
holds the hex relative-virtual-address (RVA), keyed by (lib_index, rva) so the same hex address in two different libs is never merged.
The script:
- Auto-detects the symbolicator (
--symbolicator auto):- macOS →
atos -o <path> -arch <arch> -l <base> <addrs>; the main executable's__TEXTbase is0x100000000, shared dylibs use base 0. - Linux →
addr2line -fC -e <path>; PIE binaries store RVAs directly in the profile (no base offset to add). The script reshapes the output to<func> (in <lib>) (<file:line>)so the(in <binary>)filter works on both platforms.
- macOS →
- Weights samples by
threadCPUDelta(real CPU), not wall-clockweight— otherwise parked tokio/rayon worker threads bury everything. - Reports five views:
- CPU by library (self) — where the leaf frames land. A high
libc.so.6 / libsystem_*ratio means your hot path is in syscalls/memcpy, not Rust code. - Top SELF — all — flat hot list with foreign code mixed in.
Useful to spot allocator / hashing /
memcpypressure. - Top SELF / INCLUSIVE —
wadoonly — the Rust-only view. INCLUSIVE is deduped per sample so recursive frames don't push percentages above 100%. - Syscall/alloc CPU attributed to nearest Rust caller — walks up
each non-
wadoleaf stack until it finds awadoframe and credits the cost there. This is how you find which Rust function is responsible for the__memcpy/mmap/ mimalloc hot spots. - Allocation cost by requesting caller — every sample crossing an
allocator entry, credited above the outermost one, skipping the
Vec/HashMapgrowth plumbing. The header is allocation's share of total CPU, so "is this allocation-bound at all" is a number, not a guess.
- CPU by library (self) — where the leaf frames land. A high
Common invocations:
# Default: top 30, auto-symbolicator
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts /tmp/prof.json
# Wider view; force Linux symbolicator even on macOS
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts \
/tmp/prof.json --top 60 --symbolicator addr2line
# Profile a different binary
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts \
/tmp/prof.json --binary wado-lspMemory
Measure peak RSS first, per phase second, per allocation site last.
Per phase: the span trace
--log-level debug prints every compiler span with the process's resident set
(Linux only). An end line adds the net change in current RSS since the span
began:
wado compile --log-level debug hello.wado 2>&1 | grep '<< '
# [00:00:02.1547] << stdlib_snapshot · rss 343/343 MiB (+225)The pair is current/peak MiB. A span that allocates and frees again nets out
near +0, so read a jump in the peak as well as the change. RSS is
process-wide, so under wado test with more than one worker the changes mix
every worker's compile. Use -p 1 there.
Per allocation site: DHAT
A heap profiler sees only the system allocator. wado-cli's default
mimalloc feature replaces it, so build without it, and copy the binary aside
so the next build does not replace it:
cargo build -p wado-cli --bin wado --no-default-features
cp target/debug/wado /tmp/wado-sysalloc
valgrind --tool=dhat --num-callers=40 --dhat-out-file=/tmp/dhat.json \
/tmp/wado-sysalloc compile -O2 hello.wado -o /tmp/out.wasm
node .claude/skills/profiling-wado-compiler/scripts/analyze_dhat.ts /tmp/dhat.jsonDHAT runs about 10× slower than native. Its default of 12 frames cuts off the
recursive phases, which is why --num-callers=40 is there.
The analyzer reports what was live at the heap's peak (t-gmax), since that
sets peak memory, by allocation site and by wado function inclusive.
--where RE and --not RE keep or drop stacks, so
--where get_or_init_snapshot splits the per-thread stdlib snapshot from the
compile itself. --stacks N prints the largest stacks whole.
Non-obvious points
- Read CPU, not wall-clock. The script weights by
threadCPUDelta; otherwise parked tokio/rayon worker threads bury everything. - Kernel syscall names from
atosare wrong (shared-cache base offset). Read syscall cost via the script's "nearest Rust caller" attribution, not the syscall name. - A generic hot spot may be dozens of call sites wearing one name.
Identical monomorphizations are folded to a single address, and the
symbolicator reports whichever name sorts first. One pass then appears
to own cost that belongs to the whole compiler. A dev build folds 125
Body::for_each_child::<closure>instantiations onto one address, and it reads asoptimize::drveat ~10 % self CPU while thenir/drvespan measures 0.8 %. A span that contradicts the profile is the tell. Before crediting a pass, runnm target/debug/wado | grep <symbol>and count how many names share the address. addr2lineoutermost-frame only.-i(inlined frames) is intentionally omitted because addr2line does not emit a per-input separator with-i, so the output cannot be reliably split back to addresses. The outermost frame matches the self-CPU bucket — which is what you want. If you need inlined frames, usesamply load.- Symbolication needs the matching binary. A saved profile holds only
addresses; the script resolves them against the binary at the recorded
path, so rebuilding
target/profiling/wado(ortarget/debug/wado) makes earlier profiles re-symbolicate to garbage. Analyze before rebuilding, or keep the matching binary. - Validate with the profile, not req/s or wall time. The CPU breakdown is reproducible run to run; throughput on a busy dev machine swings by tens of percent. Use the profile to confirm a change landed (e.g. a hot function shrank); measure absolute throughput on a quiet/target host.
- Linux profile includes the ld-linux frames. Unwinder occasionally
attributes a frame to
ld-linux-x86-64.so.2when the RVA happens to collide between the main binary and the dynamic linker mapping. The library breakdown's self-CPU view (which uses the leaf frame's lib) is the trustworthy ratio — incl-by-lib can over-count those frames.