K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

Local AI Hardware Benchmarks, Edge AI Agent Performance, and Sandbox Updates

Published , 04:44 Bogota (UTC-5) · 24 sourced items, 17 new since the previous edition · Read the foundations review · RSS

Today in 5 points

Laptop AI (MacBook, MLX)

SiliconBench: Speed, Memory, and Fidelity for LLM Serving on Unified-Memory DesktopsNEW

arXiv · · Laptop AI (MacBook, MLX)

SiliconBench evaluates nine Apple Silicon serving engines for concurrent local LLM serving, focusing on speed, memory headroom, and output fidelity. It uses Qwen3, Qwen3.5, and Gemma 4 models.

Why it matters: This directly addresses local AI hardware, comparing Apple Silicon performance with references like DGX Spark and evaluating specific models and runtimes like vllm-metal.

Towards Training Private LLMs: Exploring Fine-Tuning Language Models on Apple Silicon with RDMA over Thunderbolt

arXiv · · Laptop AI (MacBook, MLX)

This paper studies Apple Silicon as a platform for private LLM fine-tuning, characterizing RDMA-over-Thunderbolt (TB) communication on Mac Studio nodes. Measured bandwidth was below nominal TB specifications.

Why it matters: This is directly relevant to local AI hardware, specifically Mac Studio, for private LLM fine-tuning, highlighting both potential and current limitations of distributed execution.

Phone and edge AI

Radio-Frequency Convolutional Neural NetworksNEW

arXiv · · Phone and edge AI

RF-CNNs repurpose existing communication hardware, specifically frequency mixers in wireless radios, to perform CNN inference on edge devices. This offers an alternative to adding dedicated computing hardware.

Why it matters: This research explores using existing device hardware for AI inference, potentially reducing the need for additional NPUs/TPUs on phones and other edge devices.

Sample Count Is Not Enough: Candidate-Generation Strategy Shapes the Energy and Performance of LLM Test-Time ScalingNEW

arXiv · · Phone and edge AI

Test-time scaling improves LLM reasoning by generating multiple candidate responses. Increasing the number of candidates from 1 to 8 improved accuracy for Phi-3-mini and Qwen2.5-1.5B on GSM8K prompts.

Why it matters: This shows how inference strategies impact LLM performance on phones, highlighting the trade-offs between accuracy and system cost for models like Phi-3-mini and Qwen2.5-1.5B.

Scene-Conditioned Relation Routing for urban cellular activity forecastingNEW

arXiv · · Phone and edge AI

SCRR-Net is a scene-conditioned spatial relation routing framework for urban cellular activity forecasting. It jointly models heterogeneous spatiotemporal signals and adapts to changing urban scenes.

Why it matters: This work addresses complex AI tasks on edge devices by proposing a framework for forecasting urban cellular activity, relevant for phone-based applications.

Calibrated RF-Fingerprinting Under Interference With Heterogeneous Transmission ProtocolsNEW

arXiv · · Phone and edge AI

This work advances RF-Fingerprinting, a spectrum monitoring technique, by considering co-channel interference with multiple overlapping signals. It uses a 1D convolutional neural network for multi-label classification.

Why it matters: This research explores using CNNs for signal processing on devices, which could be relevant for phone AI applications involving radio frequency analysis.

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

NVIDIA Technical Blog · · Phone and edge AI

TensorRT Edge-LLM completed the MLPerf Edge Agentic Benchmark 6.4x faster on Jetson AGX Thor. This indicates AI agents are moving to edge devices like vehicles and robots.

Why it matters: This is a direct example of phone/edge AI performance, showing significant speedup for agentic benchmarks on an edge NPU/TPU like Jetson AGX Thor.

v0.17.1

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.1 was released, introducing a bug fix to a tool call integer type issue.

Why it matters: This update for LiteRT-LM, likely an edge or mobile runtime, addresses a bug that could affect tool calling functionality for phone AI.

Claude Cowork and chat are now one Claude

Simon Willison · · Phone and edge AI

Claude Cowork and chat are merging into a single Claude experience, which is described as becoming a general agent. This change is rolling out to web, desktop, and mobile platforms.

Why it matters: This indicates a trend towards unified agent experiences across various devices, including mobile, relevant for phone AI and agent sandboxes.

Runtimes and quantization

QUALS: Corpus Equilibrium for Universal Forecasting via Pattern Quantization and Learnability SynchronizationNEW

arXiv · · Runtimes and quantization

QUALS is a large-scale time series corpus equilibrium framework designed to enhance data efficiency for training foundation models for zero-shot forecasting. It allows existing models to achieve superior performance with a smaller fraction of training data.

Why it matters: Improved data efficiency in training can lead to more performant models that are potentially smaller or require less data, impacting inference runtimes and resource use.

Score Centering Stabilizes Off-policy Reinforcement LearningNEW

arXiv · · Runtimes and quantization

Score centering is an additive correction term that stabilizes reinforcement learning (RL) under training-inference mismatch (TIM). It addresses drift, a persistent bias between training and inference engines.

Why it matters: Stabilizing RL under TIM is crucial for deploying LLMs, especially when fine-tuning or running agents where training and inference environments may differ.

PrefixBench-H100: Characterizing Prefix Reuse and Time-to-First-Token in H100 LLM ServingNEW

arXiv · · Runtimes and quantization

PrefixBench-H100 is a benchmark and measurement framework for characterizing prefix reuse and time-to-first-token in LLM serving on a single NVIDIA H100. It evaluates vLLM and TensorRT-LLM.

Why it matters: This benchmark provides insights into optimizing LLM serving performance on high-end hardware, relevant for understanding the capabilities of OEM boxes and DGX Spark.

VLN on the Fly: An Onboard Vision-Language Navigation Stack for Aerial RobotsNEW

arXiv · · Runtimes and quantization

VLN on the Fly is an onboard vision-language navigation stack for aerial robots, keeping grounding, planning, and control as separate stages. It uses a quantized VLM for instruction grounding.

Why it matters: This demonstrates an onboard AI stack for complex tasks, showing how quantized VLMs and modular design can enable AI on resource-constrained edge devices.

v0.34.2NEW

Ollama releases · · Runtimes and quantization

Ollama v0.34.2 was released, adding a first-run setup and fixing excessive memory growth during long generations with MLX speculative decoding. It also updated llama.cpp.

Why it matters: This update improves the stability and efficiency of Ollama, a popular runtime for local LLM inference, especially for users on Apple Silicon (MLX).

b11028NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp release b11028 lists supported platforms, including macOS Apple Silicon, macOS Intel, various Linux configurations (CPU, Vulkan, CUDA, ROCm, OpenVINO, SYCL), Android arm64, and Windows.

Why it matters: This shows the broad platform support for llama.cpp, a core runtime for local LLM inference, covering a wide range of local AI hardware.

Where the Industry Is Investing: A Look at MLPerf Inference v6.1NEW

MLCommons · · Runtimes and quantization

MLPerf Inference v6.1 saw a record field of 30 submitters and 120 systems, with the introduction of agentic and end-to-end benchmarks.

Why it matters: The inclusion of agentic benchmarks in MLPerf Inference is significant for evaluating the performance of AI agents, which are often run in sandboxes or locally.

Agent sandboxes (E2B and peers)

How Lark Uses E2B to Safely Test Apps with Customer Data

E2B Blog · · Agent sandboxes (E2B and peers)

Lark uses E2B sandboxes to safely test customer applications within Docker-based development environments.

Why it matters: This provides a real-world example of E2B sandboxes being used for secure application testing, relevant for agent development and deployment.

e2b@2.50.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK e2b@2.50.0 was released, removing V1 template build operations and schemas from the generated API clients. Template builds now go through the Template SDK.

Why it matters: This update streamlines the E2B SDK, impacting how users interact with and build templates for agent sandboxes.

Open models for local use

CliniCIRCA: A Modular LLM Framework for Constructing Longitudinal Mental Health Patient Journeys from Raw EHR NarrativesNEW

arXiv · · Open models for local use

CliniCIRCA is a multi-stage LLM framework designed to reconstruct longitudinal mental health patient journeys from unstructured EHR narratives. It temporally classifies clinical events without event-level timestamps.

Why it matters: This demonstrates an application of LLMs for complex data extraction and temporal reasoning, which could inform future agentic or local LLM development.

Chain-of-Thought Entropy as a Reliability Signal: A Preregistered ReproductionNEW

arXiv · · Open models for local use

An empirical study reproduced findings that the shape of an LLM's chain-of-thought entropy trajectory predicts final answer correctness. It used four open-weight models on GSM8K and MATH-500 benchmarks.

Why it matters: Understanding LLM reliability signals is important for developing more robust AI agents and applications, whether run locally or in sandboxes.

Quantifying Overclaiming Propensity in Frontier LLM AgentsNEW

arXiv · · Open models for local use

OverclaimBench is an evaluation suite that quantifies the propensity of frontier LLM agents to overclaim task completion. It defines overclaiming as a final response contradicting information in its context.

Why it matters: This benchmark is crucial for assessing the reliability and trustworthiness of AI agents, which are often run in sandboxes or locally.

V\={a}kQA: A Benchmark and Evaluation Study for Telugu Spoken Factoid Question AnsweringNEW

arXiv · · Open models for local use

VākQA is a Telugu Spoken Question Answering (SQA) benchmark with 2,001 factoid question-answer pairs and speech audio. It evaluates proprietary and open-weight models.

Why it matters: This benchmark contributes to multilingual LLM development, which can broaden the applicability of local and phone AI models.

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive