Building general agents, in the open.
We publish the research behind Macrodeep's agents — the learning algorithms, world models, evaluation, and oversight that turn a language model into a system you can hand a goal and trust to pursue it.
GTM-Bench — Designing an Evaluation for Go-to-Market–Competent Language Models
The language-model field measures what it can grade — and there is no rigorous evaluation for the operations that actually define go-to-market work. This design note puts an axis under that capability: seven task categories, five deterministically graded against a hidden oracle, and an execution-based ORCHESTRATE category run inside a mock-MCP sandbox. It reports no scores; the leaderboard is a separate release.
Read the paper →Inside ORCHESTRATE — Grading Agentic Tool-Use in GTM-Bench
How an agent's multi-step tool use is scored against a deterministic end state — the mock-MCP sandbox, the assertion language, and the failure modes execution grading catches that prose grading waves through.
Keeping a Benchmark Honest — Contamination, Held-Out Splits, and Decay
The machinery that keeps a number from meaning less than it claims: three distinct contamination threats, a private authoritative split, and a contamination policy published as a first-class artifact.
What “Correct” Means — Constructing Ground Truth for the Objective Categories
Under the harness and the leaderboard sits the oracle. How ground truth is built for ENRICH, QUALIFY, TAM, and ROUTE — canonical record, realized outcome, curated set, decision structure — and what makes each trustworthy enough to grade against.
The GTM-Bench Harness — Architecture of a Reproducible Evaluation Runner
An engineering companion to the series: the four-stage pipeline, the thin model adapter, pure-function graders, and the containerized reproducibility guarantee that lets someone outside run the harness and get our number.
Grading the Ungradable — Rubric Methodology for the Generative Categories
COMPOSE and SUMMARIZE have no deterministic oracle. How they're graded anyway — pairwise comparison against a held-out reference, rubric-bound axes, and a judge held separate from any model under test — and why they're kept to a minority of the score.
Agency
Long-horizon decision-making — planning, recovery from error, and finishing multi-step work in open-ended environments.
Generality
Skills that transfer across tasks, tools, and domains, so one agent carries what it learns from one problem to the next.
Oversight
Evaluation and control that scale with capability, keeping agents legible and steerable as they grow stronger.
Foundations
World models, credit assignment, and the learning algorithms that let agents imagine consequences before acting.
Do research that
acts on the world.
We're hiring research scientists and research engineers who want to publish frontier work and see it ship into production.