Research

Building general agents, in the open.

We publish the research behind Macrodeep's agents — the learning algorithms, world models, evaluation, and oversight that turn a language model into a system you can hand a goal and trust to pursue it.

replace · figure / diagramFEATURED
Evaluation · Featured

GTM-Bench — Designing an Evaluation for Go-to-Market–Competent Language Models

The language-model field measures what it can grade — and there is no rigorous evaluation for the operations that actually define go-to-market work. This design note puts an axis under that capability: seven task categories, five deterministically graded against a hidden oracle, and an execution-based ORCHESTRATE category run inside a mock-MCP sandbox. It reports no scores; the leaderboard is a separate release.

Read the paper
Publications
Evaluation

Inside ORCHESTRATE — Grading Agentic Tool-Use in GTM-Bench

How an agent's multi-step tool use is scored against a deterministic end state — the mock-MCP sandbox, the assertion language, and the failure modes execution grading catches that prose grading waves through.

2026Methodology14 min
Evaluation

Keeping a Benchmark Honest — Contamination, Held-Out Splits, and Decay

The machinery that keeps a number from meaning less than it claims: three distinct contamination threats, a private authoritative split, and a contamination policy published as a first-class artifact.

2026Methodology13 min
Evaluation

What “Correct” Means — Constructing Ground Truth for the Objective Categories

Under the harness and the leaderboard sits the oracle. How ground truth is built for ENRICH, QUALIFY, TAM, and ROUTE — canonical record, realized outcome, curated set, decision structure — and what makes each trustworthy enough to grade against.

2026Methodology15 min
Evaluation

The GTM-Bench Harness — Architecture of a Reproducible Evaluation Runner

An engineering companion to the series: the four-stage pipeline, the thin model adapter, pure-function graders, and the containerized reproducibility guarantee that lets someone outside run the harness and get our number.

2026Engineering16 min
Evaluation

Grading the Ungradable — Rubric Methodology for the Generative Categories

COMPOSE and SUMMARIZE have no deterministic oracle. How they're graded anyway — pairwise comparison against a held-out reference, rubric-bound axes, and a judge held separate from any model under test — and why they're kept to a minority of the score.

2026Methodology13 min
Areas of focus
01

Agency

Long-horizon decision-making — planning, recovery from error, and finishing multi-step work in open-ended environments.

02

Generality

Skills that transfer across tasks, tools, and domains, so one agent carries what it learns from one problem to the next.

03

Oversight

Evaluation and control that scale with capability, keeping agents legible and steerable as they grow stronger.

04

Foundations

World models, credit assignment, and the learning algorithms that let agents imagine consequences before acting.

Do research that
acts on the world.

We're hiring research scientists and research engineers who want to publish frontier work and see it ship into production.