Skip to content
All work

Agentic Data Science System

LLM planning over deterministic analytics, with full traceability and evaluation on top of MCP tool execution.

PythonMCPMulti-LLM

Architecture

QuestionLLM plannerMCP toolsDeterministicanalyticsStructuredoutputTrace store: every plan, call, resulttraceability + evaluationreplayable

The problem

Letting an LLM 'do data science' directly is unreliable: it hallucinates results, can't be reproduced, and leaves no audit trail. But the planning and interpretation an LLM offers is genuinely useful. The two need to be separated.

What I built

  • A planner that decomposes an analytical question into a sequence of tool calls.
  • An MCP tool layer wrapping deterministic analytics. The LLM never computes; it orchestrates.
  • Full traceability: every plan, tool invocation and result is recorded.
  • An evaluation layer scoring whether the system reached correct, well-grounded conclusions.
  • Multi-LLM support so the planner isn't tied to one provider.

Why it matters

It keeps the LLM where it's strong (planning, interpretation) and deterministic code where it's strong (computation), giving you analyses that are both flexible and reproducible.

System notes

  • Deterministic tools mean identical inputs always produce identical outputs; the LLM's variability is confined to planning.
  • The trace is the artifact: you can replay exactly how a conclusion was reached.
  • MCP gives the tools a stable contract the planner can target across model providers.

Key decisions

LLM plans, tools compute
Never let the model produce numbers directly. Determinism lives in code; judgement lives in the planner.
MCP as the tool boundary
A standard tool protocol decouples the analytics from any single LLM and makes the system extensible.
Evaluation over trajectories, not just answers
Scoring the path, not only the final number, catches reasoning that's right for the wrong reasons.