Agents Need Fake Companies: Build a Simulated Company for AI Agent Testing

Era by Eon launched fake companies Oct 5; its benchmark grades by code, not AI judges. Write the fixture spec agents sit before a model swap: seed, graders.

Hero illustration for a simulated company for AI agent testing: a small fake company drawn as connected system tiles (CRM, helpdesk, chat, docs) around a seed value, with an answer key locked beside itHero illustration for a simulated company for AI agent testing: a small fake company drawn as connected system tiles (CRM, helpdesk, chat, docs) around a seed value, with an answer key locked beside it
A fake company is only useful if the answer key was written before the agent arrived.

Every agent tested in Eon’s September benchmark, combined, got one answer right in 84 attempts on its hardest questions: pick the right customer when the records hold two who look almost the same. Those agents were not working in anyone’s real CRM. They were working in a fake company where every answer was already known, which is the only reason anyone can say exactly how often they were wrong.

That is the case for a simulated company for AI agent testing. Before you swap the model under a support agent, widen its tool scope or upgrade its harness, it should sit the same exam in the same fake company, graded by code that knows the answers. Production cannot give you that, because production does not know its own ground truth. A staging copy cannot either, because nobody wrote down what the right answer was.

This piece gives you the exam’s spec: the entities, the seeded truth, the systems the agent sees, what it may write, how each answer is graded, how the world resets, what a run costs and who owns it. It is filled in for an illustrative support-agent eval, with the grader code and a buy-vs-build decision at the end.

Era by Eon, Oct 5: what the console, llms.txt and two papers say

On Oct 5, 2026, Eon, a cloud-backup company, opened Era to builders, according to StartupHub.ai’s launch report, which credits Eon co-founder and CEO Ofir Ehrlich. Era generates a synthetic company from a few answers (industry, business model, size, scenario, connectors) or one CLI line such as era new --industry fintech --size mid --systems salesforce,zendesk,slack. The same seeded cast of employees, customers and deals then appears across emulated Salesforce, HubSpot, Zendesk, Gong, Jira, Confluence, Slack, SharePoint, Google Drive, Box, Dropbox, AWS and Deel, with “additional connectors coming soon” per the console (read Oct 8, 2026), each serving a vendor-style REST API and an MCP endpoint.

Era by Eon console landing page reading “Complete simulated companies for your agents to run against,” with Generate a company and Free for builders buttons, a Request access link, and a note that systems are exposed through MCP servers and vendor-compatible APIs Screenshot: Era by Eon, “Era by Eon - super realistic simulated companies” (console home page, undated; launched Oct 5, 2026), captured Oct 7, 2026.

Three launch claims need their fine print. **"Free for builders"** is the console's phrase; Eon's machine-readable [llms.txt](https://console.era.eon.io/llms.txt) adds that access is granted per account on request and that usage is metered, with each API request or MCP tool call counting as one. No paid tiers were published as of Oct 8.

Writes are in conflict. The console describes a write made over MCP showing up for a REST client, as if changes propagate across systems. The llms.txt says the hosted data plane refuses writes with 403 writes_disabled, and that a product that has to write back into the vendor “cannot be exercised end to end here.” Both pages still said so when re-read Oct 8, 2026. Until Eon reconciles the two, treat hosted Era as read-only.

“Exact grading” is a property of Eon’s benchmark, not a console feature. The Sep 9 paper says “Every expected answer is computed from the final records, so grading is exact”: code computes each answer from the seeded data, with no AI judge and no grading of the agent’s trajectory. Eon reports nine models scoring 42.4% to 76.8% on 33 questions, three attempts each, and its own abstract adds that only three of 36 pairwise differences between models held up after statistical correction. The Sep 24 follow-up tested 12 agents on hidden-knowledge questions: the best got 18 of 24, and the near-duplicate questions produced that 1 in 84.

arXiv abstract page for 2609.09853, The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents, with the abstract stating every expected answer is computed from the final records Screenshot: arXiv, “[2609.09853] The Era by Eon Benchmark: A Generated Enterprise Estate with Exact Ground Truth for Benchmarking LLM Agents” (submitted Sep 9, 2026), captured Oct 7, 2026.

Both papers are Eon-authored, and so are the realism scores in them. Eon’s own docs say Era is not a substitute for a staging system with real data: the emulators implement the interfaces clients use, not complete vendor APIs.

Era is also not the first fake company. τ-bench (Jun 2024) graded retail and airline agents on the database state they left behind. TheAgentCompany (Dec 2024) built a small simulated software firm where the best agent finished 30% of tasks on its own. CRMArena-Pro (May 2025) found agents near zero on confidentiality awareness in a CRM.

τ²-bench (Jun 2025) added a telecom domain where the simulated customer also acts on the system. Snorkel’s simulated companies (Aug 3, 2026) split grading into answer, resulting state, policy violations and process. What Era adds is a generated estate wide enough to cover a whole company’s SaaS footprint.

Why the staging copy and the LLM judge both fail this test

A staging copy of your CRM has real-looking data and no answer key. When the agent says three enterprise customers have open P1 tickets, you cannot check that without doing the work yourself, so in practice nobody checks. An LLM judge fills the gap with an opinion, and its opinion drifts when its own model changes. Neither tells you whether last week’s agent and this week’s agent differ.

A seeded fixture flips the order: you generate the world, so you know every fact in it before the run, and code can compare the agent’s answer and the agent’s changes against what must be true. Checkpoint grading on end state works for computer-use runs for the same reason. The market story of who sells these environments is told in the RL environments piece; this is the operator’s side.

Write the synthetic-company fixture spec

Four rules come before any table:

  1. Exact graders before any LLM judge. If code cannot compute the expected answer from the seed, the task is not ready. A judge may score tone in a side column; it never decides pass or fail.
  2. One fixture per agent role. The support agent’s fake company is not the finance agent’s. Shared fixtures grow until nobody knows which tasks matter to which lane.
  3. Reset between runs. Every run starts from the same snapshot, or run two grades run one’s leftovers.
  4. Log the fixture version with every eval result. A score without its fixture version is not comparable to anything.

Step 1: Fill in the fixture spec, one row per entity

The example below is an illustrative support-agent eval: a mid-size online retailer whose agent answers order questions, issues refunds within policy, tags tickets and escalates to the right engineer. Counts, costs and owners are illustrative.

Entity (count) Seeded ground truth, stored as Systems exposed Write policy Exact grader (what it computes) Reset procedure Cost per run (illustrative) Owner
Customers (1,200, incl. 40 near-duplicate pairs) Seed graph; customers.parquet + SHA-256 in the manifest; duplicates listed in mess-manifest.json CRM, mail Read-only The customer ID the task refers to, resolved from order number and email, not name Restore day1 snapshot ~6 calls Support ops lead
Orders (8,400) Seeded orders table; keys/orders.json with refund-eligible flags precomputed Orders REST API Sandboxed writes (refund endpoint only, self-run simulator) Refund amount = policy formula applied to order total and days since delivery, on the fixture’s “today” Drop and reload from seed ~8 calls Billing engineer
Tickets (3,100) Helpdesk emulator export; expected end-state diff per task in keys/tickets/ Helpdesk, chat Sandboxed writes (status and tags only) The ticket diff after the run equals the expected diff, and no other ticket changed Restore day1 snapshot ~7 calls Support ops lead
Employees (85) HR seed; on-call calendar keyed to the fixture date Chat directory, docs None Escalation target = the on-call engineer for the ticket’s product on the fixture’s “today” Static; verify hash ~2 calls Platform on-call owner
Policies (14 docs) Markdown with version hashes and effective dates Docs store None Cited policy ID and version = the policy in force on the order date Static; verify hash ~2 calls Support policy owner

That is roughly 25 metered calls per task. The table is the fixture; the rest of this section is how to keep it honest.

Two columns get skipped most often. The owner column names a person who answers when a grader looks wrong, and it should be whoever owns the real policy, not whoever wrote the eval: the billing engineer knows when the refund formula changed. The cost column is there because a hosted estate meters calls and your own fixture burns compute, and either way an agent that loops on search can quietly turn a cheap exam into an expensive one. Write both before the first run, while nobody has a score to defend.

Step 2: Pin the world in a manifest

The exam must be identical every time it runs. Era’s docs name the knobs it exposes (SIMCORE_SEED, SIMCORE_TODAY and an image tag); a self-built fixture needs the same three plus its mess list. An illustrative manifest:

fixture: support-agent-eval
version: 2026.10.07-a
seed: 41873
today: "2026-09-30"
images:
  crm: crm-sim:1.4.2
  helpdesk: helpdesk-sim:0.9.0
day_state: day1
mess:
  near_duplicate_customers: 40
  conflicting_addresses: 25
  manifest: mess-manifest.json
writes: sandboxed
reset: restore_snapshot day1

Seed the mess on purpose and keep the manifest. Eon’s 1-in-84 result came from near-duplicate records, which is exactly the mistake a support agent makes when it refunds the wrong Jane Doe.

Step 3: Write the exact graders as code

Each grader is a short function that reads the seed, computes the expected answer, and compares it with what the agent produced or changed. Grade the answer and the resulting state separately. An illustrative refund grader:

def expected_refund(order, policy, today):
    days = (today - order.delivered_on).days
    if days > policy.window_days:
        return 0.0
    return round(order.total * policy.refund_share, 2)

def grade_refund(task, seed_db, final_db, today):
    order = seed_db.orders[task.order_id]
    want = expected_refund(order, seed_db.policies["REFUND-3"], today)
    got = final_db.refunds.get(task.order_id, 0.0)
    stray = diff_tables(seed_db, final_db) - {("refunds", task.order_id)}
    return {"answer_ok": got == want, "state_ok": not stray}

The stray check is the one teams skip. An agent that issues the right refund and also closes three unrelated tickets has a right answer and a wrong state, and a right answer never offsets a wrong state. Score them in separate columns, add a third for permission violations, and fail the task if any column fails.

Step 4: Keep state-change tasks off the hosted plane

Lookup tasks run anywhere. Tasks that change records need a fixture that accepts writes and a grader that diffs records before and after, the way τ-bench graded the database its agents left behind. If you use hosted Era, its llms.txt says writes return 403, so run those tasks on self-run simulators or your own seeded database. Keep the ratio honest: a fixture with only lookup questions tests reading, and the support agent’s riskiest moves are writes.

Step 5: Run each task three times and pass on all three

One lucky run proves little. Sierra’s τ-bench post introduced pass^k, the chance that an agent succeeds on all k tries, and showed GPT-4o falling to about 25% at pass^8 in its retail domain. Use k=3 as a floor: a task passes only if all three runs pass every column. Eon’s three-attempt design follows the same instinct.

Step 6: Reset to the snapshot between every run

Restore the day1 snapshot (or your equivalent) before each run and check its hash. Era’s day-states are a usable model: day0 empty, day1 with full history, day2 set 91 days later for tasks that depend on time passing. A fixture without a reset step quietly turns run three into a test of whether the agent can clean up after run two.

Step 7: Log every result with its fixture version

Every eval result is one line, keyed to the fixture version and seed, so a later comparison can refuse to mix versions. An illustrative record:

{"fixture": "support-agent-eval", "fixture_version": "2026.10.07-a", "seed": 41873,
 "lane": "support-triage", "model": "candidate-b", "harness": "3.2.0", "k": 3,
 "pass_k": "31/40", "answer_fail": 4, "state_fail": 5, "violations": 0,
 "metered_calls": 3010, "date": "2026-10-07", "approver": "support-ops-lead"}

That line is what a CI regression gate reads when it decides whether the change ships; the gate’s thresholds live there, not here. If a skill is heading for a schedule, this fixture is also its test bed before it climbs the ladder to a cron.

Loop diagram of a synthetic-company fixture: seed, fixture, agent run, exact grader, result log and reset back to the seed snapshotLoop diagram of a synthetic-company fixture: seed, fixture, agent run, exact grader, result log and reset back to the seed snapshot Six stops and a loop. The reset arrow is the one that keeps every run the same exam.

The worked scenario: one model swap, one exam

All numbers illustrative. The support agent above is moving from its current model to a candidate. The exam has 40 tasks: 28 lookups and 12 state changes. At k=3 that is 120 runs, at about 25 metered calls each, roughly 3,000 calls per candidate.

Model tokens come to about 60K per run, 7.2M per candidate, or $21.60 at an illustrative blended $3 per million.

The current model passes 33 of 40 at pass^3 with zero violations. The candidate passes 31 of 40, fails five tasks on state (four of them refunds to the near-duplicate customer) and has zero violations. Its lookup scores are higher than the current model’s, and a best-of-three average would have called it an upgrade.

The state column says it refunds the wrong person more often, so it does not ship until a prompt or tool change fixes the duplicate resolution and the exam is re-run on the same fixture version.

Timeline of dated simulated-company and agent-environment benchmarks: tau-bench June 2024, TheAgentCompany December 2024, CRMArena-Pro May 2025, tau-squared-bench June 2025, Snorkel’s simulated companies August 2026, Era benchmark paper September 9, 2026, Era launch October 5, 2026Timeline of dated simulated-company and agent-environment benchmarks: tau-bench June 2024, TheAgentCompany December 2024, CRMArena-Pro May 2025, tau-squared-bench June 2025, Snorkel’s simulated companies August 2026, Era benchmark paper September 9, 2026, Era launch October 5, 2026 Seven dated releases from the papers and posts cited here. Era is the latest entry in a line that started in 2024, not the start of it.

Buy the hosted estate or build your own seeded fixture

Question Hosted generated estate (Era-style) Your own seeded fixture
Time to a first exam Hours, once access is approved Days to weeks
State-change tasks Not on hosted Era today; self-run simulators needed Yes, you control writes
Fidelity to your stack Emulated vendor interfaces, not your config or custom fields Matches your schema and policies
Ground truth for graders Seeded data available; you still write the graders Yours by construction
Cost shape Metered per call, free tier gated Engineering time plus hosting
Version control Seed, date and image tag pinnable Whatever you build
Data sensitivity No real data leaves you No real data, if you seed from scratch

Buy when you need breadth fast and your tasks are mostly lookups across many SaaS systems. Build when the agent’s job is writes against your own schema, or when your policies are the thing under test. Many teams will do both: a hosted estate for broad read exams, a small self-built fixture for the five write paths that can hurt a customer.

How fixture-based agent evals go wrong

What breaks Signal you would see First action
Fixture drift Scores jump on an unchanged agent; the result log shows two fixture versions in one comparison Re-run the baseline on the pinned version; block cross-version comparisons
Answer key computed from the wrong state Grader fails runs whose transcript shows the correct action Recompute keys from final records; hand-check two failed runs
Reset skipped Pass rate differs by run index (run one passes, run three fails) Hash-check the snapshot before every run
Agent learns the fixture Fixture scores climb while production escalations do not fall Regenerate with a new seed each quarter; keep the old seed as a regression set
Hosted write refusal read as agent failure Every state-change task errors with 403 writes_disabled Move write tasks to a self-run simulator
Judge creep Pass or fail flips when the judge model is updated Move judge scores to a side column that never gates
Metered calls run away Call count per run spikes as an agent loops on search Cap calls per task; treat hitting the cap as a failure

One exam for every lane in the fleet

A fleet with several agent CLIs and several models has no fair way to compare them except a shared, fixed exam. The fixture is that exam, and its result log is the record of which lane passed which version on which date with whose approval. That makes it part of the operating layer for agents, next to the schedules and permissions it informs, rather than a side project owned by whoever ran the last benchmark.

The fixture tests whether an agent does its job correctly. Whether it holds up when someone attacks it through a tool or a poisoned record is a separate plan, covered in the MCP integration attack checklist. If you need the basics of what an agent eval is in the first place, the evals explainer covers that ground.

FAQ

What is a simulated company for AI agent testing?

A generated business with fake customers, orders, tickets, employees and policies spread across emulated systems such as a CRM, helpdesk and chat. Because the data is seeded, every correct answer is known in advance, so code can grade an agent’s answers and changes exactly instead of relying on a human or an AI judge.

Is Era by Eon free, and can agents write to it?

Era’s console says “Free for builders,” but its llms.txt says access is approved per account and every API request or MCP tool call is metered. The llms.txt also says the hosted version refuses writes with a 403 error, while the console implies writes propagate. Check both before planning write tests.

How is a synthetic business environment different from tau-bench?

τ-bench, released by Sierra in June 2024, tests agents in narrower retail and airline domains and grades the database state they leave behind. Generated estates such as Era cover many SaaS systems with one shared cast of people. Both rely on known ground truth; the newer ones trade depth per domain for breadth.

Sources

YOU'RE THROUGH THIS ONE.

Keep connecting the dots.

Back to the library