
Oqoqo: Evals and private benchmarks for real-world agent tasks
Renzo Viale and Haritha Nair co-founded Oqoqo so teams can measure whether agents can actually use their product, on the workflows that matter to them.
Oqoqo turns a team’s real workflows into reproducible evals and private benchmarks; companies building for agents learn where agents succeed, where they stall, and what fixing that is worth.
Oqoqo starts with a real task: the project state, files, data, tools, and services an agent needs, plus a rubric that defines what success means. It runs that task across models, agents, and product versions, each trial in its own sandbox, then reports pass rates, token spend, and every step the agent took to finish or fail.
Public benchmarks measure general model capability, which is the right question for the labs and the wrong one for a team shipping a product. Agents do not operate in general; they operate inside a company’s workflows, data, tools, and interfaces. Trying an agent once is an anecdote. Oqoqo repeats the same tasks under the same conditions, so the comparison holds.
The user at the top of the developer funnel is shifting from a human evaluating a product page to an agent reading docs and running calls. That shift breaks assumptions baked into how developer-tool companies have written, tested, and marketed their products for two decades. Oqoqo gives those companies the measurement layer that shift requires: evals they own, benchmarks built on their own success criteria, and regression suites that catch failures as models and products change.
The new user doesn’t leave feedback; it’s a coding agent.
Co-Founders: Renzo Viale & Haritha Nair
Last Funding Stage: Pre-Seed
$1.5m
Funding to Date


