AI Setup Foundations: Local AI and E2B
Updated 2026-10-09 · daily research brief at 04:44 Bogota · today's brief
A working review of the two halves of an agent stack: Local AI (where the model thinks, on hardware you own) and E2B (where the agent acts, inside a disposable sandbox). Categories first, then the setups, then how they fit together.
01
ai_setup_foundations: shared categories
Every setup, local or sandboxed, is judged on the same eight categories. Amber = Local AI concern, cyan = E2B concern.
1. Compute
LocalGPU/unified memory capacity sets which models fit; memory bandwidth sets tokens/sec.
E2BvCPU and RAM per sandbox; no GPU inference inside.
2. Runtime
LocalMLX, llama.cpp, Ollama, LM Studio, vLLM.
E2BFirecracker microVM with a Linux userland, driven by SDK.
3. Models
LocalOpen-weight, quantized (Q4 to Q8, MLX 4/8-bit). MoE preferred.
E2BModel-agnostic: any LLM (local or API) issues the commands.
4. Isolation
LocalProcess level only; the model runs with your user's reach.
E2BKernel-level VM boundary per session; the reason E2B exists.
5. Data residency
LocalPrompts and weights never leave the box. PII-safe.
E2BCloud by default; code and files transit their infra unless self-hosted.
6. Interfaces
LocalOpenAI-compatible HTTP on localhost.
E2BPython/JS SDK, MCP server, Code Interpreter, Desktop (computer use).
7. Cost model
LocalCapEx up front, near-zero marginal cost per token.
E2BOpEx per sandbox-second; free tier, then usage billing.
8. Operations
LocalYou patch, cool, update models, monitor thermals.
E2BManaged lifecycle; you own templates, timeouts, cleanup.
02
Local AI setups reviewed
Five hardware archetypes. The constraint that matters most is how many GB the model can live in, not raw compute.
| Setup | Model memory | Strengths | Weaknesses | Best for |
|---|---|---|---|---|
| Apple Silicon workstation Mac Studio M5 Max / Ultra | 96 to 256 GB unified | Large models in one address space, silent, low watts, MLX native, no driver work | Slower prefill than NVIDIA; no CUDA; vLLM ecosystem limited | Single operator running 30B to 120B MoE agents, private data |
| Apple Silicon laptop MacBook Air/Pro 16 to 48 GB | ~10 to 36 GB usable | Portable, zero setup | Only small models (3B to 14B, small MoE); swap kills speed | Autocomplete, drafts, offline fallback |
| Consumer NVIDIA rig 1 to 4x RTX 4090/5090 | 24 to 32 GB per card | Fastest tok/s per dollar, CUDA, vLLM/SGLang, fine-tuning | VRAM split across cards, power (450 to 600 W each), noise, heat | Throughput, batching, LoRA training, serving a team |
| Pro NVIDIA workstation RTX PRO 6000 96 GB, DGX Spark class | 96 to 128 GB | Big models on CUDA in one device, datacenter tooling | Highest price per GB; Spark-class boxes have modest bandwidth | CUDA-first shops, research, fine-tuning 70B class |
| AMD unified APU Ryzen AI Max (Strix Halo) mini PCs | up to ~96 GB of 128 GB | Cheapest route to large memory, Linux friendly | ROCm/Vulkan maturity, lower bandwidth, fewer tuned kernels | Budget large-model experimentation |
| Rented GPU (hybrid) Vast.ai, RunPod, H100/H200 | 80 to 141 GB per GPU | No CapEx, burst to any size | Data leaves premises, hourly cost compounds, cold starts | Occasional heavy jobs; never PII |
Sizing rule: largest model file + 15 to 20 GB for KV cache and OS = minimum memory tier. A 84 GB model therefore needs the 128 GB tier.
03
Runtimes
MLX / mlx-lm
AppleApple's array framework. Fastest path on Apple Silicon, native quantization, OpenAI-style server via mlx_lm.server.
llama.cpp
UniversalC/C++ engine, GGUF format. Runs on Metal, CUDA, ROCm, Vulkan, CPU. The portable fallback and the engine under many GUIs.
Ollama
ConvenienceOne-command pull and serve, model registry, localhost API on :11434. Trades tuning control for simplicity.
LM Studio
GUIDesktop app with MLX and llama.cpp backends, model browser, local server. Good for evaluating quants side by side.
vLLM / SGLang
NVIDIA servingPagedAttention, continuous batching, tensor parallel. The choice when many requests hit one model.
Open WebUI / agents
Front endChat UI, RAG and tool calling on top of any OpenAI-compatible endpoint; agent frameworks point here too.
04
Models and sizing
Mixture-of-experts dominates local: large total capacity, 3B to 12B active parameters per token, so speed stays usable.
| Model | Quantized size | Role | Minimum tier |
|---|---|---|---|
| Qwen3-Coder 30B | ~18 GB | Fast coder | 36 GB |
| Qwen3.6-35B-A3B | ~33 GB | Daily coder | 64 GB |
| Nemotron 3.5 Lightning 30B | ~35 GB | Agent driver (Hermes) | 64 GB |
| Qwen3.5-122B-A10B | ~64 GB | Long context (262k) | 96 GB |
| gpt-oss-120B | ~65 GB | General reasoning | 96 GB |
| Nemotron 3 Super 120B | ~84 GB | Heavy reasoning | 128 GB |
Sizes are approximate for 4-bit class quants; long contexts add KV cache on top.
05
E2B: sandboxes for agents
E2B is open-source infrastructure that gives an AI agent a secure, disposable Linux computer. The LLM decides what to run; E2B runs it somewhere that cannot touch your host.
Isolation
FirecrackerEach sandbox is a microVM (the tech behind AWS Lambda). Starts in well under a second; its own kernel boundary.
SDKs
Python / JSSandbox.create(), run commands, read/write files, stream stdout, expose ports, kill on exit.
Code Interpreter
Jupyter kernelStateful cells, charts returned as data. The data-analysis agent pattern.
Desktop
Computer useSandbox with a graphical desktop and stream; agent clicks and types inside the VM, not on your machine.
Templates
Custom imagesBake dependencies into a template so each session starts ready; persistence and pause/resume for long jobs.
Deployment
Cloud or self-hostManaged cloud with usage billing and a free tier; open-source infra can be self-hosted (Terraform, GCP/AWS) when data must stay in-house.
# the canonical loop: model proposes, sandbox executes
from e2b_code_interpreter import Sandbox
from openai import OpenAI # pointed at a LOCAL endpoint
llm = OpenAI(base_url="http://localhost:1234/v1", api_key="local")
with Sandbox.create() as sbx:
code = llm.chat.completions.create(model="qwen3-coder-30b",
messages=[{"role":"user","content":"Load data.csv and plot revenue by week"}]
).choices[0].message.content
result = sbx.run_code(code) # runs in the microVM, not on the host
print(result.logs, result.error)
When to use E2B: any agent that writes and runs code, installs packages, browses, or operates a desktop. When not: pure chat, retrieval, or classification with no execution.
06
Sandbox options compared
| Option | Boundary | Startup | GPU | Data stays local | Fit |
|---|---|---|---|---|---|
| E2B cloud | Firecracker microVM | sub-second | No | No | Fastest route to safe agent execution |
| E2B self-hosted | Firecracker microVM | sub-second | No | Yes (your cloud) | Same API, compliance-bound data |
| Docker on the AI box | Container (shared kernel) | ~1 s | Possible | Yes | Trusted code, lowest cost, weaker isolation |
| gVisor / Kata locally | User-space kernel / light VM | 1 to 2 s | Limited | Yes | Stronger local isolation, more ops work |
| Daytona / Modal | Managed containers / VMs | sub-second to seconds | Modal: yes | No | Dev environments, GPU functions |
| No sandbox | None | 0 | n/a | Yes | Never for model-written code on a production host |
07
Reference stack: local brain, sandboxed hands
Tier A: private
Local model + Docker/gVisor on the same box. Nothing leaves. Use for anything with PII.
Tier B: hybrid
Local model + E2B cloud for execution on non-sensitive code and public data. Best safety per hour of setup.
Tier C: compliance
Local model + self-hosted E2B in own VPC. Strong isolation and residency; highest ops load.
Rule: prompts with client PII never reach a cloud model or cloud sandbox. Route by data class, not by convenience.
08
Verdict
Local AI: Mac Studio M5 Max 128 GB is the floor for a six-model MoE lineup up to 84 GB; MLX first, llama.cpp as fallback, Ollama/LM Studio for convenience. Move to NVIDIA only when throughput or fine-tuning becomes the bottleneck.
E2B: adopt as the execution layer for any agent that runs code. Start on the cloud tier with synthetic or public data; switch to self-hosted or a local gVisor sandbox before any production data enters the loop.
Together: the model thinks locally, the agent acts in a box. Capacity decides the hardware, data class decides the sandbox.
09
Deep dives: three folders
Click a folder to open it; click it again to close.
16 GB unified leaves roughly 10 to 11 GB for weights plus KV cache after macOS. Target models of 8B or less at 4-bit, or small MoE. Memory bandwidth sets speed: M4 ~120 GB/s, M5 ~150 GB/s (M5 also adds GPU neural accelerators that speed up prefill in MLX).
| Model (4-bit) | Size | M4 Air 16 GB | M5 Pro-chassis 16 GB | Use |
|---|---|---|---|---|
| Qwen3 4B / Qwen3.5 4B class | ~2.5 GB | ~40 to 55 tok/s | ~50 to 65 tok/s | Fast drafts, tool calling |
| Phi-4-mini 3.8B | ~2.3 GB | ~45 tok/s | ~55 tok/s | Reasoning, math, short context |
| Gemma 3 / Gemma 4 small (4B to E4B) | ~3 GB | ~35 tok/s | ~45 tok/s | Multilingual ES/EN, vision |
| Qwen3 8B / Llama 3.1 8B | ~4.5 to 5 GB | ~20 to 25 tok/s | ~25 to 30 tok/s | Best quality that fits comfortably |
| Qwen3-Coder 30B-A3B (3-bit) | ~13 GB | swaps, unusable | borderline, close all apps | Not recommended at 16 GB |
Throughput figures are indicative ranges for MLX 4-bit at short context; verify on your own box with mlx_lm.generate --verbose.
MLX (first choice)
pip install mlx-lm mlx_lm.server --model mlx-community/Qwen3-8B-4bit --port 1234
Similar runtimes
- LM Studio: GUI, MLX + GGUF backends
- Ollama: one-command, llama.cpp under the hood
- llama.cpp: Metal, GGUF, most quant options
- Apple Foundation Models: ~3B on-device model via Swift, free, private
16 GB survival rules
- Raise GPU wired limit:
sudo sysctl iogpu.wired_limit_mb=12288(resets on reboot) - Keep context at 8k to 16k; KV cache grows fast
- Close browsers before loading; watch memory pressure, not free RAM
- Air throttles on long runs (fanless); Pro chassis sustains
Verdict
Good for autocomplete, private drafts, offline fallback and an agent that calls a bigger remote brain. Not a replacement for the 128 GB box. If buying new: 24 to 32 GB is the real laptop floor.
| Box | Chip | Memory | Bandwidth | Approx. price | Strength | Weakness |
|---|---|---|---|---|---|---|
| Mac Studio M5 Max | Apple M5 Max | up to 128 GB | ~550+ GB/s | ~$5.4k (128 GB) | Fast decode, silent, MLX, macOS | No CUDA, slower prefill than NVIDIA, no clustering story |
| Mac Studio M3/M5 Ultra | Apple Ultra | 96 to 512 GB | ~800 GB/s | $5.5k to $10k+ | Largest single-box models (200B to 600B+ quantized) | Price, prefill on long prompts |
| NVIDIA DGX Spark | GB10 Grace Blackwell | 128 GB | 273 GB/s | ~$4k | Full CUDA stack, fast prefill, FP4, fine-tuning, 200 GbE pair-up (2 units = 405B class) | Decode ~2x slower than Max/Ultra, ARM Linux (DGX OS) |
| ASUS Ascent GX10 | GB10 | 128 GB | 273 GB/s | ~$3k to $3.5k | Same silicon as Spark, cheaper, smaller SSD tiers | Same as Spark |
| Dell Pro Max GB10 / HP ZGX Nano / Lenovo ThinkStation PGX / Acer Veriton GN100 / MSI EdgeXpert | GB10 | 128 GB | 273 GB/s | ~$3k to $4k | Enterprise support, procurement friendly | Differentiation is chassis, SSD and warranty only |
| AMD Strix Halo (Framework Desktop, GMKtec EVO-X2, HP Z2 Mini G1a) | Ryzen AI Max+ 395 | 128 GB (~96 to 110 GB to GPU) | 256 GB/s | ~$2k to $2.5k | Cheapest 128 GB, x86, Linux or Windows | ROCm/Vulkan maturity, slowest prefill of the group |
| RTX PRO 6000 workstation | Blackwell 96 GB | 96 GB VRAM | 1.8 TB/s | ~$9k+ card alone | Fastest by far, vLLM serving | Price, 600 W, 96 GB cap per card |
Decode (tokens out)
Bound by bandwidth. Mac Max/Ultra win: roughly 2x Spark on the same MoE model.
Prefill (reading prompts)
Bound by compute. Spark/GB10 wins, often several times faster on long contexts and RAG, where agents spend most time.
Pick Mac Studio if
One operator, chat and coding agents, silence, macOS workflow, MLX.
Pick GB10 (Spark or OEM) if
CUDA code must run as-is, fine-tuning/LoRA, long-context agent prefill, or prototyping for DGX cloud. Buy the OEM unit unless NVIDIA support matters.
Prices and bandwidth are approximate public figures as of Sep 2026; verify in configurators.
Here E2B means Gemma's "effective 2B" edge variant (per-layer embeddings, ~2B active footprint), not the E2B sandbox above. Phones run 1B to 4B models at 4-bit; RAM above 12 GB is what separates usable from toy.
| Model | Size (4-bit) | Strength | Runtime path |
|---|---|---|---|
| Gemma 3n / Gemma 4 E2B | ~1.5 to 2 GB | Text, image, audio in; built for mobile | Google AI Edge Gallery, LiteRT-LM, MediaPipe LLM |
| Gemma E4B | ~3 GB | Better quality, still phone class | Same; needs 12 GB+ phone |
| Qwen3 0.6B / 1.7B / 4B | 0.5 to 2.5 GB | Tool calling, multilingual, thinking mode | PocketPal (llama.cpp), MLC Chat, MNN Chat |
| Phi-4-mini 3.8B | ~2.3 GB | Reasoning and math per byte | PocketPal, ONNX Runtime GenAI, Termux llama.cpp |
| Gemini Nano | system | OS-integrated, zero setup | AICore on Pixel and select Samsung |
| Phone | Chip / accelerator | RAM | Notes |
|---|---|---|---|
| Pixel 10 Pro / Pro XL | Tensor G5 (TSMC 3 nm) + Google TPU | 16 GB | Best Gemma and Gemini Nano integration; TPU used via AICore/LiteRT, GPU path for open models |
| Galaxy S25/S26 Ultra | Snapdragon 8 Elite class, Hexagon NPU | 12 to 16 GB | Fastest raw CPU/GPU decode on Android; Qualcomm AI Hub NPU models |
| OnePlus 13/15, ROG Phone, Xiaomi 15/17 Ultra | Snapdragon 8 Elite class | 16 to 24 GB | Most RAM for the money; room for 7B to 8B at 4-bit |
| iPhone 17 Pro | A19 Pro + GPU neural accelerators | 12 GB | MLX Swift apps (Locally AI, PocketPal iOS), Apple Foundation Models |
| Galaxy S22 (existing) | Snapdragon 8 Gen 1 / Exynos 2200 | 8 GB | Qwen3 1.7B or Gemma E2B only; thermals limit long runs |
Setup A: zero config
Install Google AI Edge Gallery, download Gemma E2B, run offline chat, image Q&A and audio transcription.
Setup B: any GGUF
PocketPal AI (Android/iOS): pull Qwen3 4B or Phi-4-mini Q4_K_M from Hugging Face, set context 4k.
Setup C: phone as server
pkg install llama-cpp # Termux llama-server -m qwen3-4b-q4_k_m.gguf \ --host 127.0.0.1 --port 8080 -c 4096
Localhost OpenAI API for scripts on the phone; tunnel only with auth.
Expectations
- E2B / 1.7B: ~15 to 30 tok/s on flagship
- 4B: ~8 to 15 tok/s
- Thermal throttle after minutes; battery drain high
- Best use: offline, private triage, not agents
10
Daily research brief: 2026-10-09
Updated every night by 04:44 Bogota from arXiv, Apple, NVIDIA, Google, Microsoft Research and official release feeds. Last edition .
Local AI Runtimes Advance, Phone AI Gets Multimodal, and Agent Sandboxes Enhance Security
- Ollama v0.40.2 upgrades models for improved performance and llama.cpp compatibility, while llama.cpp release b11514 includes a Musa FWHT fix and broad platform support. NVIDIA DGX Spark is becoming available with 64GB of unified memory, supporting local AI development, and experiments on a local DGX
- Google AI Edge Gallery 1.0.20 and LiteRT-LM v0.18.0 both feature official support for EmbeddingGemma 2, enabling on-device multimodal semantic search with NPU and GPU acceleration. The Apache 2.0 license for EmbeddingGemma 2 is noted for its benefits in applications requiring many embedding vectors.
- E2B sandboxes are used by ClickUp for running AI code on sensitive data with isolated microVMs, and E2B Secrets are introduced for safe credential injection. OpenAI "rogue" agent activities were found on Wikimedia projects, highlighting security challenges. Cowork has shifted its model inference and
- b11516: hexagon: fix IM2COL patch-embed DMA ring overflow (#30189) llama.cpp releases · 2026-10-09
- Self-Organization from Constrained Geometric Radiation arXiv · 2026-10-08
- Coverage-Aware Reasoning with Medical Tokens for Diagnosis Prediction arXiv · 2026-10-08
- KDFP: A first-principles approach to knowledge distillation in large language models arXiv · 2026-10-08
- Adaptive Multi-Discriminator WGAN Framework for Resource-Constrained Internet of Vehicles Using Reinforcement Learning and Game Theory arXiv · 2026-10-08
- Rethinking the Tradeoff Between Temporal Encoding and Nonlinear Computation in Spiking Language Models arXiv · 2026-10-08
11
Frequently asked questions
Can a MacBook Air M4 with 16 GB RAM run a local LLM?
Yes, models up to about 8B parameters at 4-bit (roughly 5 GB) run well with MLX or llama.cpp. Qwen3 8B runs around 20 to 25 tokens per second on an M4; 30B-class models swap and are not practical at 16 GB.
What is the best runtime for local AI on Apple Silicon?
MLX (mlx-lm) is usually the fastest on Apple Silicon. LM Studio offers MLX and GGUF backends in a GUI, Ollama is the simplest command-line option, and llama.cpp is the portable fallback.
Mac Studio or NVIDIA DGX Spark for local AI?
Mac Studio M5 Max or Ultra generates tokens faster because of higher memory bandwidth. DGX Spark and GB10 OEM boxes process long prompts faster and run the full CUDA stack, which matters for fine-tuning and CUDA-only code.
What are the DGX Spark OEM alternatives?
ASUS Ascent GX10, Dell Pro Max with GB10, HP ZGX Nano, Lenovo ThinkStation PGX, Acer Veriton GN100 and MSI EdgeXpert use the same GB10 Grace Blackwell chip with 128 GB unified memory; AMD Strix Halo mini PCs are the lower-cost non-NVIDIA alternative.
How much memory do I need to run a 120B model locally?
Take the quantized model file size and add 15 to 20 GB for KV cache and the OS. An 84 GB 120B-class quant needs the 128 GB tier; a 65 GB model fits in 96 GB with short context.
Which phones run local AI models best?
Phones with 12 to 16 GB RAM and strong NPUs: Pixel 10 Pro (Tensor G5 with Google TPU), Galaxy S25/S26 Ultra and other Snapdragon 8 Elite phones, and iPhone 17 Pro. They run 1B to 4B models such as Gemma E2B, Qwen3 4B and Phi-4-mini.
What is Gemma E2B?
Gemma E2B is Google's effective-2B edge variant of Gemma, designed for phones and laptops, with text, image and audio input. It runs offline through Google AI Edge Gallery, LiteRT-LM and MediaPipe. It is unrelated to the E2B sandbox company.
What is E2B and why use it with AI agents?
E2B is open-source infrastructure that gives AI agents secure, disposable Linux sandboxes built on Firecracker microVMs. The agent runs generated code inside the sandbox instead of on your machine, via Python or JavaScript SDKs.
Can E2B be used with a local LLM?
Yes. E2B is model-agnostic: a local model served by MLX, Ollama or LM Studio on an OpenAI-compatible endpoint decides what code to run, and the E2B sandbox executes it. Self-host E2B when data must not leave your infrastructure.
Is local AI cheaper than cloud GPUs?
For daily use it usually is. Renting a 96 GB GPU at about 1 to 2 USD per hour for 4 hours a day breaks even with a 128 GB workstation in roughly 9 to 34 months, and local keeps private data on premises.