WorxAI / marketplace

Flagship dataset / agent action benchmarks

Ship agents with data you can replay.

A privacy-safe synthetic corpus for evaluating tool choice, checkout behavior, latency, retries, and recovery before production. Start with a free preview, then scale through API delivery or monthly refreshes.

99.1%

replay consistency

Deterministic schemas for agent evaluation and regression testing.

10k+

benchmark rows

A ready-to-run sample pack for tool choice, checkout, and recovery traces.

Base USDC

agent-native delivery

Pay per row through x402 or choose a monthly refresh plan.

Free operator sample

For operators / watch first

See the render before you buy.

Watch a free synthetic motion sample, then inspect the generator route, provenance manifest, and delivery options. No checkout is required to view it.

What is included

A practical evaluation layer for autonomous systems

  • XITool-selection traces with success and failure outcomes
  • XICheckout and payment-flow scenarios with bounded retries
  • XILatency, recovery, and escalation fields for SLO testing
  • XIStable provenance metadata and synthetic disclosure on every row
  • XIFree preview plus API, bulk, and monthly refresh delivery

Machine contract

Built for agents, not spreadsheets

View plans →
GET /api/v1/data/agent-action-benchmarks/preview
GET /api/x402/v2/datasets/detail?id=agent-action-benchmarks
POST /api/x402/v2/answers