Episode Endpoints — Simulation & Agent Evaluation
Three MCP tools for retrieving simulation episode data — agent logs, replays, and per-submission episode listings. Useful when evaluating agent-based hackathon submissions or auditing simulation runs.
list_submission_episodes — ✅ PASS
List the episodes generated by a competition submission.
Parameters:
submissionId(int) — the Kaggle submission id
Use this as the entry point: pull the episode ids first, then fetch logs or replays per episode.
get_episode_agent_logs — 🔬 BAD_PROBE
Download per-agent log output for one episode.
Parameters:
episodeId(int)agentIndex(int) — 0-based agent index within the episode
Status: Marked BAD_PROBE in the 2026-04-22 audit (test infrastructure issue rather than confirmed tool failure). The tool is documented and likely works; verify behavior before depending on it. Capture both success and error shapes when you call it.
from skills.kaggle.shared.mcp_client import mcp_call, resolve_token
token = resolve_token()
resp = mcp_call("get_episode_agent_logs", {
"episodeId": 12345678,
"agentIndex": 0,
}, token=token)get_episode_replay — 🔬 BAD_PROBE
Download the replay payload for one episode.
Parameters:
episodeId(int)
Status: Same caveat as get_episode_agent_logs. Tool exists; full live
behavior unverified. Treat the response shape as discoverable rather than
documented until a clean probe lands.
Recommended pattern
# 1. Pull the submission's episodes
episodes = mcp_call("list_submission_episodes", {"submissionId": SID}, token=token)
# 2. For each episode, optionally pull replay + per-agent logs
for ep_id in extracted_episode_ids:
replay = mcp_call("get_episode_replay", {"episodeId": ep_id}, token=token)
for i in range(num_agents):
logs = mcp_call("get_episode_agent_logs", {
"episodeId": ep_id, "agentIndex": i,
}, token=token)When the BAD_PROBE endpoints return unexpected shapes, surface the raw response to the caller — a future audit may reclassify them.