AgBench Released: Evaluating Agentic AI Workloads on Personal Devices
AgBench introduces a systematic benchmark suite for assessing agentic AI performance on local, hybrid, and cloud architectures, providing reproducible results and comparative insights for developers deploying on personal AI devices.
AgBench is a new open-source benchmark suite designed to evaluate agentic AI workloads across different deployment architectures—including local, hybrid, and cloud—specifically for personal AI devices. With this release, developers can systematically measure task success rates, latency, cloud API cost, and data exposure for their agentic systems.

Experimental results using AgBench, with data from over 162 million samples, show that while many agent tasks can be executed locally on personal devices, local-only execution typically has lower task success rates and longer completion times compared to cloud-based execution, especially under high concurrency. However, running tasks locally removes cloud API costs and prevents sensitive data exposure to cloud models.
Hybrid execution strategies—splitting tasks between local and cloud resources—can improve task success but introduce variable cloud costs and potential data exposure depending on how tasks are partitioned. Importantly, no single architecture delivered optimal results across all metrics. The choice between local, hybrid, or cloud deployment should be based on application workload and device capabilities.
AgBench and all benchmark artifacts are openly available at https://anonymous.4open.science/r/AgBench-2777.
