Code labs
Broken and half-finished code drawn from mistakes that are easy to make and hard to spot. Diagnose the problem, then compare against a working solution.
17 labs
Labs are free to work through without an account. Sign in to have your progress remembered.
D1Agentic Architecture & Orchestration
- Parallel tool resultsAn agent that queries three services has started making tool calls one at a time instead of in parallel, roughly a week …fix
- Losing tool_use blocks in historyThis loop fails on the second iteration with an error about a tool_use id having no matching result — even though the co…fix
- Completing the agentic loopFill in the stop-reason handling for a manual tool-use loop that also uses a server-side tool, so it terminates correctl…fill
- Breaking out of a Managed Agents sessionA Managed Agents client returns as soon as the session reports idle. It works for simple prompts, but any run involving …fix
D2Tool Design & MCP Integration
- A strict schema that isn'tThis tool was marked strict so its inputs would always validate, but malformed arguments still reach the handler and the…fix
- Half an MCP configurationA developer connected a remote MCP server so the agent could file tickets. The request is rejected as invalid, and the e…fix
- A tool description that under-triggersA support agent has a search tool over the internal knowledge base, but it answers from memory far too often and only se…fill
- An unsafe text editor handlerThis handler implements the text editor tool. It passes tests and works in demos. Review it as though the model output w…fix
D3Claude Code Configuration & Workflows
- An instruction that should be a hookA team added this to CLAUDE.md to keep the codebase formatted. It works most of the time, then quietly doesn't during lo…fix
- Claude Code in CI that never finishesA nightly job runs Claude Code to triage failing tests. It produces no output and is killed by the pipeline timeout ever…fix
- A skill that never triggersThis skill was written to enforce a team's API conventions. It is installed correctly, but Claude almost never reads it …fill
D4Prompt Engineering & Structured Output
- Forcing JSON the old wayAn extraction pipeline written against an older model has started returning 400s after a model upgrade. It worked unchan…fix
- Thinking configuration after an upgradeThis request errors on a current model. It was written when a fixed thinking budget was the way to control reasoning dep…fix
- Batch results that silently mismatchA nightly job classifies 20,000 documents through the Batch API. It runs without errors, but roughly a fifth of the stor…fix
D5Context Management & Reliability
- A cache that never hitsA support assistant sends a large shared system prompt on every request. Caching was enabled weeks ago, but `cache_read_…fix
- Compaction state thrown awayA long-running assistant enables compaction so conversations can outlive the context window. It works for a while, then …fix
- Crashing on a successful responseA production service throws an IndexError a few times a day. The traceback points at response handling, the HTTP status …fix