Tool Call Bench

Staff Writer

Priya Subramaniam

Priya covers the science and politics of AI evaluation, with a background in machine learning research and a decade spent stress-testing model outputs at two Bay Area AI labs. She has a particular eye for what benchmark leaderboards reveal—and conceal—about real-world agent behavior.

1 story