About getEvals
About getEvals
The independent evaluator of artificial intelligence. We build the benchmarks and evaluation infrastructure that measure whether models can do the work of lawyers, bankers, engineers, and doctors.
Powered by PumaAI
The Measurement Gap
Rapid progress in AI has been driven by a process researchers call hill-climbing: repeatedly defining a measurable objective, identifying shortcomings, and improving against them. Every major advance in AI, from image classification to question-answering to software development, has been measured this way. One hill at a time.
The pace of model development is outpacing the community’s ability to construct new hills. Trillions have been invested in generating intelligence, but comparatively little in measuring it. As a result, models summit old benchmarks in months. Test sets released openly are absorbed into pre-training corpora, which quietly invalidates the results built on them. And new benchmarks are often produced or run by the same companies that build the models, reported alongside cherry-picked examples and evaluation regimens tuned to score well.
AI has quickly become a trillion-dollar market without the independent measurement institutions that other markets of this scale, like finance and healthcare, require. The industry is worse off because of it.
- Labs lack a credible way to demonstrate continued model progress.
- Enterprises increasingly see pressure to adopt and spend on AI without means to quantify ROI.
- Governments must develop their own expertise to measure frontier capabilities, cyber risk, and the pace of global competition.
We started getEvals to solve this problem, as the independent evaluator of artificial intelligence, powered by PumaAI.
What We Build
We build benchmarks that measure the ability of models to do the work of lawyers, bankers, engineers, and doctors. This takes partnering with reference institutions in each field to build a taxonomy of representative tasks, and developing new methodologies to automatically score the quality of generated work product.
Scores are based on our privately held test sets to preserve the integrity and signal of our results. This prevents training to the test set, unlike in open-source benchmarks.
We have built the infrastructure we rely on to run these evaluations reproducibly and at scale across labs. This includes a distributed system to run agentic benchmarks, and a model library — a standard API to call models with.
We also produce benchmarks we believe will be broadly beneficial to the public good. For example, Public Benefits Bench measures whether AI can be trusted to answer SNAP questions for the 37 million families who rely on the program.
Get in touch
Whether you build AI, evaluate it, or want to help shape a benchmark in your field, we would love to hear from you.
Or send us an email at contact@getevals.com