Ai News
Ai News

Liquid AI and Artificial Analysis Launch Pipette for On-Device AI Benchmarking

Published Aug 24, 2026 Sources checked Aug 27, 2026

Pipette is an open-source benchmarking platform for edge AI deployments, launching with 1,000+ model, quantization, runtime, device and context configurations plus reproducibility protocols.

Pipette benchmarks the deployed system, not just the model card

Liquid AI and Artificial Analysis released Pipette on August 24, 2026, an open-source benchmarking platform designed to measure how foundation models actually behave on edge devices.

The project starts from a practical observation: a model’s production behavior on a phone or laptop depends on more than its parameter count or full-precision quality score. Quantization, runtime, processor, context length, memory pressure and thermal state can all change whether a model is usable.

Pipette therefore records deployment configurations as combinations of model + quantization + runtime + device, then measures performance under versioned benchmark definitions and published protocols.

The launch dataset covers more than 1,000 deployment configurations

Liquid AI says the initial public dataset contains five on-device performance metrics across more than 1,000 model × quantization × runtime × device × context configurations, spanning more than 30 models. Context lengths range from 256 to 8,192 tokens.

Initial device coverage includes a MacBook Pro with M5 Max, iPhone 17 Pro and Galaxy S26 Ultra. Liquid AI says AMD Ryzen AI Max+ 395 and Radeon 8060S results are planned. The benchmark clients support macOS, Windows, iOS and Android, including native iOS and Android applications for running tests on target hardware.

Community-submitted result publication is currently in beta.

Pipette separates quality from speed, latency and memory

The interactive dashboard connects quality evaluations with measured time to first token, end-to-end latency, prefill throughput, decode throughput and peak memory. Quality and performance remain separate dimensions rather than being collapsed into a single score.

That separation matters for edge deployment. A smaller or more heavily quantized model may respond faster but lose important task quality. A sparse model may activate relatively few parameters per token yet still require memory for all of its expert weights. Two models with similar size can also scale very differently as context length increases.

Pipette’s quality evaluations currently include IFBench, GPQA Diamond and MATH-500, with compatible full-precision references where available. Liquid AI says the current quality completions are generated on NVIDIA H100 80GB reference systems and then paired with compatible on-device performance measurements.

Reproducibility is a first-class part of the release

Liquid AI says timing and memory benchmarks use fixed token shapes, greedy decoding, a discarded warm-up, five measured repetitions and readiness checks before each run. Platform-specific checks verify acceptable thermal and load conditions before a result is published.

The company says Artificial Analysis reviewed and verified the measurement methodology. The scoring path is separated from generation provenance so evaluation outputs can be scored deterministically without the scorer knowing which model produced them.

Three Apache 2.0-licensed repositories implement the pipeline: management and submission infrastructure, device benchmark clients, and model-blind scoring components. Results retain benchmark versions, runtime versions, device details and measurement conditions so comparisons can be audited later.

Early results illustrate why deployment context matters

Liquid AI highlights examples where configuration choices materially change conclusions. On Galaxy S26 Ultra, two 350M Granite variants retain very different proportions of decode throughput as context grows from 256 to 4,096 input tokens. The company also shows an iPhone comparison where MiniCPM5-1B finishes a defined workload faster than LFM2.5-1.2B-Instruct, while LFM scores higher on MATH-500 — a direct speed-versus-quality tradeoff.

These examples are useful demonstrations of the framework rather than universal rankings. Liquid AI itself warns against treating current cross-device results as controlled hardware comparisons because Android and iOS paths use different execution environments and accelerators.

Important limitations remain

Pipette does not yet provide broad NPU coverage. The current Android path is CPU-focused under the tested stable backends, while iPhone runs use Metal. The initial quality suite also does not comprehensively cover multimodal tasks, knowledge-intensive workloads or agentic behavior.

Those limitations are significant, but documenting them is part of the project’s value. Edge AI benchmarks are easy to misuse when model artifacts, quantization levels, runtimes and hardware conditions are mixed together. Pipette’s contribution is a public structure for keeping those variables explicit.

For teams deciding whether an open model can run acceptably on a phone, laptop or embedded system, reproducible deployment measurements can be more actionable than server-class model-card scores alone.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books