Back Issues/Search Home → Calendar → Archive → RSS → Subscribe → Current Issue → Popular →

All issues › Volume 342, Issue 5 › IT News › AI

Building an Enterprise AI Benchmark Changed How I Evaluate AI

CIO, Thursday, October 1st, 2026

Nutanix co-founder Dheeraj Pandey says enterprise AI evaluations should test context assembly and permissions, not just model intelligence.

Dheeraj Pandey, who co-founded Nutanix and now leads DevRev, draws on database benchmark history to argue that public AI leaderboards say little about enterprise performance, noting that while building an enterprise benchmark his team found the main bottleneck was assembling the right context with correct permissions, not model reasoning.

He recommends that CIOs build evaluations that reproduce their own operating conditions, use business-critical questions rather than polished demos, and require agents to prove reliable reading before granting write access.

The key question, he says, is whether a system can assemble the right context for the right person and prove that it did.

more →  ·  More from AI →