---
name: profiling-wado-compiler
description: Profile the native Rust `wado` binary (compile/serve/run) for host-side bottlenecks — CPU with a sampling profiler, memory with the span trace's RSS and valgrind DHAT. Use for native CPU or memory profiling, not guest wasm (see wado-performance for that).
---

# Profiling the native `wado` binary

Host-side Rust profiling (the compiler, `wado serve`, `wado run`, …
including wasmtime/cranelift). For the **guest** wasm program, use
`wado-performance` instead.

## Pick a build profile

Choose based on **what you're optimising for**:

| Profile     | Cargo flag                                   | Use when                                                                                                                                                                                                                                                               | Trade-off                                                                                                          |
| ----------- | -------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------ |
| `profiling` | `cargo build --profile profiling --bin wado` | Improving **benchmark scores** — release-equivalent codegen with debug info kept. Inherits `release` (thin LTO, `codegen-units=1`) and adds `debug = 2`, `strip = false`.                                                                                              | Slow build (LTO), but the CPU profile reflects what users actually run.                                            |
| `dev`       | `cargo build --bin wado`                     | Improving **developer-iteration time** — making `cargo run -- compile/test/...` faster for compiler-hackers. Uses the in-workspace `[profile.dev.package.wado-compiler] opt-level = 1`, so the compiler itself isn't molasses while everything else stays unoptimised. | Fast build, but the absolute numbers are larger than a release run; ratios between hot paths are still actionable. |

If unsure: pick `profiling` for "users complain it's slow", pick `dev`
for "rebuild → run → tweak feels slow during development." The
analyzer script and recording flow below are identical for both.

## Workflow

```sh
# 1. Build with the chosen profile (see table above)
cargo build --profile profiling --bin wado    # benchmark-oriented
# or
cargo build --bin wado                        # dev-iteration-oriented

# 2. Record under load with samply (cargo install samply)
samply record --save-only --rate 1000 -o /tmp/prof.json -- \
  target/profiling/wado serve --addr 127.0.0.1:8080 app.wado &
SAMPLY_PID=$!
# ... drive load (e.g. oha against benchmark/http_routing) ...

# 3. Stop: SIGTERM the CHILD, not samply. samply finalizes on child exit;
#    signalling samply leaves the child running and the recording hangs.
kill -TERM "$(pgrep -P "$SAMPLY_PID" | head -1)"; wait "$SAMPLY_PID"

# 4. Analyze
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts /tmp/prof.json
```

Symbols resolve against the path the profile recorded, not the binary that
produced it. Rebuilding over `target/profiling/wado` makes every earlier profile
symbolicate against the new build, and its output still looks normal. When you
A/B, give each arm its own path: copy the first build aside and profile it
there.

For one-shot commands (`wado compile foo.wado`, `wado test foo.wado`)
there is nothing to drive — samply records until the child exits, so
just invoke it directly:

```sh
samply record --save-only --rate 1000 -o /tmp/prof.json -- \
  target/debug/wado test package-gale/tests/driver_rust_test.wado
```

Interactive call tree (and correct kernel symbols):
`samply load /tmp/prof.json` opens a browser-based call-tree UI. The
CLI analyzer below is for grep-able, transcript-friendly summaries.

The analyzer is a TypeScript script run directly by Node.js (>= 23.6,
which strips types with no flags). No build step or dependencies are
needed; `node analyze_native_profile.ts ...` just works.

## Linux setup

```sh
# samply needs perf_event_paranoid <= 1 for a non-root user
echo '1' | sudo tee /proc/sys/kernel/perf_event_paranoid

# `addr2line` is part of binutils — usually already installed
addr2line --version >/dev/null || sudo apt-get install -y binutils
```

## How `analyze_native_profile.ts` works

samply's `--save-only` profile is **unsymbolicated**: `funcTable.name`
holds the hex relative-virtual-address (RVA), keyed by `(lib_index, rva)` so the same hex address in two different libs is never merged.
The script:

1. **Auto-detects the symbolicator** (`--symbolicator auto`):
   - macOS → `atos -o <path> -arch <arch> -l <base> <addrs>`; the main
     executable's `__TEXT` base is `0x100000000`, shared dylibs use base 0.
   - Linux → `addr2line -fC -e <path>`; PIE binaries store RVAs directly
     in the profile (no base offset to add). The script reshapes the
     output to `<func> (in <lib>) (<file:line>)` so the
     `(in <binary>)` filter works on both platforms.
2. **Weights samples by `threadCPUDelta`** (real CPU), not wall-clock
   `weight` — otherwise parked tokio/rayon worker threads bury everything.
3. **Reports** five views:
   - **CPU by library (self)** — where the leaf frames land. A high
     `libc.so.6 / libsystem_*` ratio means your hot path is in
     syscalls/`memcpy`, not Rust code.
   - **Top SELF — all** — flat hot list with foreign code mixed in.
     Useful to spot allocator / hashing / `memcpy` pressure.
   - **Top SELF / INCLUSIVE — `wado` only** — the Rust-only view.
     INCLUSIVE is deduped per sample so recursive frames don't push
     percentages above 100%.
   - **Syscall/alloc CPU attributed to nearest Rust caller** — walks up
     each non-`wado` leaf stack until it finds a `wado` frame and
     credits the cost there. This is how you find which Rust function
     is responsible for the `__memcpy` / `mmap` / mimalloc hot spots.
   - **Allocation cost by requesting caller** — every sample crossing an
     allocator entry, credited above the outermost one, skipping the
     `Vec`/`HashMap` growth plumbing. The header is allocation's share of
     total CPU, so "is this allocation-bound at all" is a number, not a
     guess.

Common invocations:

```sh
# Default: top 30, auto-symbolicator
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts /tmp/prof.json

# Wider view; force Linux symbolicator even on macOS
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts \
  /tmp/prof.json --top 60 --symbolicator addr2line

# Profile a different binary
node .claude/skills/profiling-wado-compiler/scripts/analyze_native_profile.ts \
  /tmp/prof.json --binary wado-lsp
```

## Memory

Measure peak RSS first, per phase second, per allocation site last.

### Per phase: the span trace

`--log-level debug` prints every compiler span with the process's resident set
(Linux only). An end line adds the net change in current RSS since the span
began:

```sh
wado compile --log-level debug hello.wado 2>&1 | grep '<< '
# [00:00:02.1547] << stdlib_snapshot · rss 343/343 MiB (+225)
```

The pair is current/peak MiB. A span that allocates and frees again nets out
near `+0`, so read a jump in the peak as well as the change. RSS is
process-wide, so under `wado test` with more than one worker the changes mix
every worker's compile. Use `-p 1` there.

### Per allocation site: DHAT

A heap profiler sees only the system allocator. `wado-cli`'s default
`mimalloc` feature replaces it, so build without it, and copy the binary aside
so the next build does not replace it:

```sh
cargo build -p wado-cli --bin wado --no-default-features
cp target/debug/wado /tmp/wado-sysalloc
valgrind --tool=dhat --num-callers=40 --dhat-out-file=/tmp/dhat.json \
  /tmp/wado-sysalloc compile -O2 hello.wado -o /tmp/out.wasm
node .claude/skills/profiling-wado-compiler/scripts/analyze_dhat.ts /tmp/dhat.json
```

DHAT runs about 10× slower than native. Its default of 12 frames cuts off the
recursive phases, which is why `--num-callers=40` is there.

The analyzer reports what was live at the heap's peak (t-gmax), since that
sets peak memory, by allocation site and by `wado` function inclusive.
`--where RE` and `--not RE` keep or drop stacks, so
`--where get_or_init_snapshot` splits the per-thread stdlib snapshot from the
compile itself. `--stacks N` prints the largest stacks whole.

## Non-obvious points

- **Read CPU, not wall-clock.** The script weights by `threadCPUDelta`;
  otherwise parked tokio/rayon worker threads bury everything.
- **Kernel syscall names from `atos` are wrong** (shared-cache base
  offset). Read syscall cost via the script's "nearest Rust caller"
  attribution, not the syscall name.
- **A generic hot spot may be dozens of call sites wearing one name.**
  Identical monomorphizations are folded to a single address, and the
  symbolicator reports whichever name sorts first. One pass then appears
  to own cost that belongs to the whole compiler. A dev build folds 125
  `Body::for_each_child::<closure>` instantiations onto one address, and
  it reads as `optimize::drve` at ~10 % self CPU while the `nir/drve`
  span measures 0.8 %. A span that contradicts the profile is the tell.
  Before crediting a pass, run `nm target/debug/wado | grep <symbol>`
  and count how many names share the address.
- **`addr2line` outermost-frame only.** `-i` (inlined frames) is
  intentionally omitted because addr2line does not emit a per-input
  separator with `-i`, so the output cannot be reliably split back to
  addresses. The outermost frame matches the self-CPU bucket — which
  is what you want. If you need inlined frames, use `samply load`.
- **Symbolication needs the matching binary.** A saved profile holds only
  addresses; the script resolves them against the binary at the recorded
  path, so rebuilding `target/profiling/wado` (or `target/debug/wado`)
  makes earlier profiles re-symbolicate to garbage. Analyze before
  rebuilding, or keep the matching binary.
- **Validate with the profile, not req/s or wall time.** The CPU
  breakdown is reproducible run to run; throughput on a busy dev
  machine swings by tens of percent. Use the profile to confirm a
  change landed (e.g. a hot function shrank); measure absolute
  throughput on a quiet/target host.
- **Linux profile includes the ld-linux frames.** Unwinder occasionally
  attributes a frame to `ld-linux-x86-64.so.2` when the RVA happens to
  collide between the main binary and the dynamic linker mapping. The
  library breakdown's self-CPU view (which uses the leaf frame's lib)
  is the trustworthy ratio — incl-by-lib can over-count those frames.
