Meet oqoqo

13 minutes ago

Oqoqo helps teams evaluate how agents use real products and workflows. Private benchmarks and repeatable experiments make it possible to compare behavior under the same tasks and see where an agent succeeds or gets stuck.

What's included in this release:

  • Teams can turn real workflows into private task sets and rubrics that stay versioned with each experiment.
  • Each trial runs in an isolated sandbox with the project state, credentials, and tools needed for a realistic session.
  • Teams can compare agents, models, effort levels, and treatments across repeated runs.
  • Full trajectories show tool calls, commands, errors, files, and stopping points so failures have useful context.
  • CLI and MCP access lets teams trigger experiments from existing development workflows, including CI changes.

Together, these capabilities create a repeatable loop for defining an evaluation, running it consistently, understanding the result, and improving what failed.