K/20X LABS · AI_SETUP_FOUNDATIONS · DAILY RESEARCH BRIEF

AI Runtimes Advance, Edge AI Performance Boosted, Agent Sandboxes Evolve

Published , 04:44 Bogota (UTC-5) · 28 sourced items, 19 new since the previous edition · Read the foundations review · RSS

Today in 5 points

K/20X research paper

Phone and edge AI

Modelling daily activity patterns from mobile phone location data via deep representation learningNEW

arXiv · · Phone and edge AI

The Activity Chain Encoder (ACE) is a self-supervised model for modeling daily activity patterns from mobile phone location data. ACE combines pre-trained urban embeddings with visit timing and duration, using a Transformer to model sequential organization of

Why it matters: This research focuses on processing mobile phone location data, indicating advancements in on-device data analysis for activity patterns, relevant for phone AI applications.

The Effect of Quantization on Clinical Benchmarks: Accuracy and Safety Across Model FamiliesNEW

arXiv · · Phone and edge AI

A study evaluated five 7-8B parameter models at FP16, GPTQ-INT8, and GPTQ-INT4 precision on clinical benchmarks. INT8 GPTQ showed universal safety, while INT4 degradation was substantial and model-dependent, with BioMistral-7B losing 19.7% on MedMCQA.

Why it matters: This research provides specific data on the impact of quantization on model accuracy and safety, especially for clinical applications on resource-constrained edge devices, informin

Efficient Mixture-of-Experts with Speculative Decoding via Expert CoactivationNEW

arXiv · · Phone and edge AI

Combining Mixture-of-Experts (MoE) models with Speculative Decoding (SD) for inference acceleration is challenging due to memory transfer costs. The study finds that MoE routers with high degrees of expert coactivation result in faster runtimes.

Why it matters: This research addresses a key bottleneck in accelerating MoE inference on NPUs, which is crucial for efficient phone AI and edge device performance.

Global Ranks Survive, Selected Heads Shift: BOS-Sink Topology under 4-bit Weight-Only QuantizationNEW

arXiv · · Phone and edge AI

A study on 4-bit NF4 weight-only post-training quantization (PTQ) found that global bf16-to-4-bit ranks for Sink-aware deployment remain high, but top-k Jaccard overlap is lower. Terminal Qwen layers show significant shifts, while Llama-3.2-1B shows low, unifo

Why it matters: This provides insights into the effects of 4-bit quantization on attention mechanisms, which is crucial for understanding and optimizing the performance of quantized models on loca

datasette-auth-github 1.0

Simon Willison · · Phone and edge AI

datasette-auth-github 1.0 fixes an issue where authenticated sessions expired prematurely on mobile Safari due to missing Max-Age parameters in cookies. The plugin is now stable and tested against Datasette 0.65.x and 1.0ax.

Why it matters: This is a fix for a web-based authentication plugin, relevant for web applications that might interact with AI agents or services, especially on mobile devices.

TensorRT Edge-LLM Completes the MLPerf Edge Agentic Benchmark 6.4x Faster on Jetson AGX Thor

NVIDIA Technical Blog · · Phone and edge AI

TensorRT Edge-LLM completed the MLPerf Edge Agentic Benchmark 6.4x faster on Jetson AGX Thor, demonstrating AI agents moving from cloud to edge devices like vehicles and robots.

Why it matters: This highlights significant performance improvements for AI agents on edge devices, directly relevant to phone AI and other NPU/TPU-equipped hardware.

v0.17.1

LiteRT-LM releases · · Phone and edge AI

LiteRT-LM v0.17.1 introduces a bug fix for a tool call integer type issue.

Why it matters: This is a minor bug fix for a runtime, indicating ongoing maintenance and refinement for efficient model execution, potentially on phone AI platforms.

Claude Cowork and chat are now one Claude

Simon Willison · · Phone and edge AI

Claude Cowork and chat are merging into one Claude, becoming a general agent. This change is rolling out to Pro and Max plans on web, desktop, and mobile.

Why it matters: This signifies a trend towards unified, general-purpose AI agents across platforms, including mobile, simplifying user experience and potentially expanding agent capabilities on ph

Runtimes and quantization

v0.30.0

vLLM releases · · Runtimes and quantization

vLLM v0.30.0 introduces 762 commits, 315 contributors, and new model support including DeepSeek-V4.1-Flash with MXFP8 KV storage on SM100, DeepGEMM Mega-mHC, and a DeepSeek-V4 CPU backend. It also features Fast Start, a persistent per-GPU weight-cache daemon f

Why it matters: This release expands model compatibility and introduces features like Fast Start and MXFP8 KV storage, which can improve inference performance and reduce restart times for local AI

b11096NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp update b11096 updates the Level Zero SDK to v1.33.1 and enables L0/oneDNN CMake.

Why it matters: This update indicates ongoing development for broader hardware support and optimization within llama.cpp, potentially benefiting local inference on various devices.

PRQuant: Permutation Residual Quantization for Low-Overhead InferenceNEW

arXiv · · Runtimes and quantization

PRQuant (Permutation Residual Quantization) is a training-free, low-overhead framework for low-bit quantization of linear layers. It combines channel reorganization with static weight-side residual compensation, constructing residual weight sub-tensors offline

Why it matters: PRQuant offers a method to improve accuracy in low-bit quantization without heavy execution overheads, relevant for efficient local inference on resource-constrained hardware.

StepKV: Step-Aware KV Cache Compression for LLM AgentsNEW

arXiv · · Runtimes and quantization

StepKV proposes step-aware KV cache compression for LLM agents, addressing the linear growth of KV cache with context length. It aims to improve upon existing token-level pruning methods by recognizing the uneven and delayed importance of reasoning steps.

Why it matters: StepKV offers a method to reduce KV cache storage and decoding costs, which is critical for efficient inference of LLM agents, especially on local or sandbox environments with limi

Towards Full Pipeline FP8 Reinforcement Learning for LLMsNEW

arXiv · · Runtimes and quantization

Full-pipeline FP8 Reinforcement Learning (RL) for LLMs suffers from training instability, manifesting as entropy surges and garbled outputs. This is traced to compounded FP8 quantization noise distorting the importance ratio.

Why it matters: This identifies a significant challenge in using FP8 quantization for RL training of LLMs, which is relevant for optimizing local training or fine-tuning of models.

Neural Residual Modeling for Scientific Data Compression under Guaranteed Error BoundsNEW

arXiv · · Runtimes and quantization

Neural Residual Modeling proposes augmenting RVQ-based compressors with a U-Net trained to predict and correct pixel-space residuals between original and RVQ-reconstructed scientific data. This addresses the spatially structured nature of residuals.

Why it matters: This research focuses on improving lossy compression for scientific data, which could have implications for efficient data handling in AI applications.

b11090NEW

llama.cpp releases · · Runtimes and quantization

llama.cpp update b11090 fixes an sm_70 tile compilation error by generalizing the tile shape of the 5-argument load_ldmatrix from <16,8> to <I,J>. This allows the non-swizzle branch to forward to the 3-argument loader for any shape.

Why it matters: This is a technical fix that improves compatibility and compilation for specific NVIDIA GPU architectures (Volta), benefiting local inference on such hardware.

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-TritonNEW

NVIDIA Technical Blog · · Runtimes and quantization

NVIDIA TensorRT multi-device inference is a new capability simplifying model serving across multiple GPUs, addressing the compute and memory demands of generative AI that exceed single GPU capacity.

Why it matters: This capability is crucial for scaling local AI inference beyond a single GPU, enabling more powerful models to run on multi-GPU setups like high-end workstations.

Agent sandboxes (E2B and peers)

A Pinch of SFT, A Dash of RL: When Reinforcement Learning Helps Long-Horizon Advertising AgentsNEW

arXiv · · Agent sandboxes (E2B and peers)

This study investigates balancing Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) for long-horizon advertising agents. It observes three regimes: Imitation, Lift, and Discovery, for improving tool-use problems.

Why it matters: This research explores methods to improve agent performance, particularly for complex, multi-step tasks, which is directly relevant to the development and training of agents in san

OSWorld-Pro: Process-based Evaluation for Computer Use AgentsNEW

arXiv · · Agent sandboxes (E2B and peers)

OSWorld-Pro is a new evaluation benchmark for Computer-Use Agents (CUAs) with over 300 tasks and 2800 subgoals, enabling procedural evaluation. It uses human-aligned LLM-Judges to assess subgoal fulfillment, providing transparency into agent failures.

Why it matters: OSWorld-Pro provides a more granular evaluation method for agents, which is crucial for understanding and improving their performance in sandbox environments.

BreakFun: Jailbreaking LLMs via Object Instantiation under Simulated Code ExecutionNEW

arXiv · · Agent sandboxes (E2B and peers)

BreakFun is a jailbreak method that frames harmful requests as code-execution simulations. It uses a "Trojan Schema" (benign Python class definition) and adversarial field names to steer the model's object instantiation towards harmful content.

Why it matters: This research identifies a new jailbreaking technique for LLMs, which is critical for enhancing the security and safety of agents operating in sandboxes.

DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at ScaleNEW

arXiv · · Agent sandboxes (E2B and peers)

DeepSeek Elastic Compute (DSec) is a production sandbox platform for large-scale agentic training and evaluation. It exposes FnCall, container, microVM, and full-VM sandbox backends through a unified SDK, coordinating placement and lifecycle management.

Why it matters: DSec provides a robust and elastic infrastructure for agentic training and evaluation, directly supporting the development and testing of agents in sandboxes.

e2b@2.51.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK v2.51.0 removes SDK-side defaults from API request payloads, allowing API defaults to apply. Sandbox create and connect now use v2 API endpoints, defaulting timeout to 5 minutes and securing envd access. The secure option on Sandbox.create is deprecate

Why it matters: This update streamlines E2B sandbox configuration by relying on API defaults and standardizes secure access, improving the developer experience for agent sandboxes.

How Lark Uses E2B to Safely Test Apps with Customer Data

E2B Blog · · Agent sandboxes (E2B and peers)

Lark uses E2B sandboxes to test customer applications in Docker-based development environments.

Why it matters: This is a real-world use case demonstrating how E2B sandboxes are used for secure application testing, reinforcing their utility for agent development and deployment.

e2b@2.50.0

E2B SDK releases · · Agent sandboxes (E2B and peers)

E2B SDK v2.50.0 removed V1 template build operations and schemas from generated API clients, as the API no longer serves them. Template builds now go through the Template SDK.

Why it matters: This is an API cleanup and migration, ensuring that E2B's sandbox template building process is streamlined and uses current methods, relevant for developers working with agent sand

Open models for local use

List Counting Failures Are Not One PhenomenonNEW

arXiv · · Open models for local use

Open-weight chat models exhibit distinct modes of failure when counting items in bracketed lists. Qwen and Gemma 27B often flip odd lengths to even, OLMo concentrates errors on mid-sized integers, and Llama tends to under-count.

Why it matters: This study reveals specific failure modes in open-weight models, which is important for understanding and improving the reliability of LLMs used locally or in sandboxes.

Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language ModelsNEW

arXiv · · Open models for local use

A study investigated how four open-weight LLMs represent clinical cost tradeoffs. Patient risk was recoverable, and cost direction was recoverable in every model. However, representational shifts tracked output changes only in larger models, and responses to c

Why it matters: This research highlights limitations in how open-weight LLMs handle complex clinical decision-making with cost tradeoffs, important for deploying such models in sensitive applicati

$t_0$: A Time-Series Foundation Model for Forecasting with ContextNEW

arXiv · · Open models for local use

$t_0$ is a family of open-weight foundation models for forecasting with multivariate context, with first members $t0$-alpha (102M parameters) and $t0$-beta (256M parameters). They use Transformer layers that alternate attention along time and across variates.

Why it matters: This introduces new open-weight models specifically designed for time-series forecasting, which could be relevant for local deployment in specialized analytical AI applications.

Does the Truthfulness Signal Survive Code-Mixing? Probing Hidden States for Hallucination Detection in HinglishNEW

arXiv · · Open models for local use

A study investigated whether hallucination probes trained on clean-language hidden states transfer to Hinglish (Hindi-English code-mixed text). It found that the signal degrades under code-mixing for three open-weight 7-8B LLMs.

Why it matters: This research highlights a challenge for hallucination detection in multilingual contexts, which is important for ensuring reliability of LLMs, especially those deployed locally or

Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

NVIDIA Technical Blog · · Open models for local use

Alibaba released Qwen3.8-2.4T-A95B (Qwen3.8-Max), its largest open-weight model with 2.4 trillion parameters, offering configurable reasoning. It can be served on NVIDIA GB300 NVL72.

Why it matters: This introduces a very large open-weight model, relevant for high-end local AI setups that can handle such scale, pushing the boundaries of what's available.

Method

Sources: arXiv API, Apple Machine Learning Research, NVIDIA, Google Research, Google Developers, Microsoft Research, Hugging Face, MLCommons, MIT News, Nature Machine Intelligence, Communications of the ACM, and official GitHub release feeds (MLX, llama.cpp, Ollama, vLLM, MLC LLM, LiteRT-LM, E2B). Items are filtered by topic rules; summaries are AI-assisted (gemini-2.5-flash) and grounded only in each source's own abstract or post text. Always read the linked source before acting.

Archive