Kamiwaza AI Launches Signal65 PINNACLE to Evaluate Real-World AI Performance

News related to:Kamiwaza AI · 2 min read

Kamiwaza AI has unveiled a new benchmarking framework, Signal65 PINNACLE, designed to evaluate the real-world performance of enterprise AI systems. The new benchmark moves beyond traditional measures to assess whether AI can successfully complete enterprise tasks, sustain correct work at scale, and do so at an acceptable cost.

According to Kamiwaza AI, the growing complexity of enterprise AI applications requires a more holistic approach to evaluation. Traditional benchmarks often focus on model capabilities or infrastructure performance in isolation, leaving businesses without a clear understanding of whether their AI systems can handle the actual work they need done. The Signal65 PINNACLE framework, co-developed with Signal65, addresses this gap by evaluating AI systems based on their ability to complete real-world enterprise workloads.

At the heart of this new benchmark is the PICARD execution harness, a technology developed by Kamiwaza AI. PICARD enables AI agents to perform multi-step jobs in generated enterprise environments, using tools like filesystems, databases, and Python runtimes. Each environment is set up with a known answer key that remains hidden from the agent, allowing the completed work to be objectively graded in code. This ensures a more accurate assessment of the AI's performance without relying on model memorization or human review.

The first PINNACLE data set tested 44 configurations across 30 base models from 12 different makers, including hosted APIs and open-weight models running on NVIDIA H200, NVIDIA B300, and AMD Instinct MI300X infrastructure. The results revealed significant variability in performance based on data conditions. Nine configurations completed at least 95% of jobs with clean, governed data, while only two configurations reached that threshold with messy data. Rankings also changed by workload, with smaller and open-weight models outperforming larger, cutting-edge models on certain tasks.

Token pricing did not reflect the cost of successful work, with input tokens representing 65% to 91% of hosted-model task costs, while output-token pricing alone understated the cost per correct task by three to 11 times. These findings underscore the importance of considering the full cost structure when evaluating AI systems.

Kamiwaza AI CEO and co-founder Luke Norris emphasized the need for flexibility in AI decision-making. "Enterprises shouldn't have to bet their AI strategy on one model or infrastructure provider, hoping it works across every workload," Norris said. "We built this technology to make it measurable, and Signal65's use of it as the foundation for PINNACLE shows how much that capability matters now."

The launch of Signal65 PINNACLE provides a new tool for enterprises to evaluate their AI models and infrastructure more accurately. Companies can use the benchmark to determine whether their current AI systems can deliver the accuracy, scale, and economics required for production. The framework offers a more nuanced and practical approach to AI evaluation, helping businesses make informed decisions about their AI strategies.

Talk to the desk

Want your company on the wire?

File your first press release free, or talk to us about a plan built for regular volume and placement.

Contact us