K/20X LABS · AI_SETUP_FOUNDATIONS

AI Setup Foundations: Local AI and E2B

Updated 2026-10-09 · daily research brief at 04:44 Bogota · today's brief

A working review of the two halves of an agent stack: Local AI (where the model thinks, on hardware you own) and E2B (where the agent acts, inside a disposable sandbox). Categories first, then the setups, then how they fit together.

01
ai_setup_foundations: shared categories

Every setup, local or sandboxed, is judged on the same eight categories. Amber = Local AI concern, cyan = E2B concern.

1. Compute

LocalGPU/unified memory capacity sets which models fit; memory bandwidth sets tokens/sec.

E2BvCPU and RAM per sandbox; no GPU inference inside.

2. Runtime

LocalMLX, llama.cpp, Ollama, LM Studio, vLLM.

E2BFirecracker microVM with a Linux userland, driven by SDK.

3. Models

LocalOpen-weight, quantized (Q4 to Q8, MLX 4/8-bit). MoE preferred.

E2BModel-agnostic: any LLM (local or API) issues the commands.

4. Isolation

LocalProcess level only; the model runs with your user's reach.

E2BKernel-level VM boundary per session; the reason E2B exists.

5. Data residency

LocalPrompts and weights never leave the box. PII-safe.

E2BCloud by default; code and files transit their infra unless self-hosted.

6. Interfaces

LocalOpenAI-compatible HTTP on localhost.

E2BPython/JS SDK, MCP server, Code Interpreter, Desktop (computer use).

7. Cost model

LocalCapEx up front, near-zero marginal cost per token.

E2BOpEx per sandbox-second; free tier, then usage billing.

8. Operations

LocalYou patch, cool, update models, monitor thermals.

E2BManaged lifecycle; you own templates, timeouts, cleanup.

02
Local AI setups reviewed

Five hardware archetypes. The constraint that matters most is how many GB the model can live in, not raw compute.

SetupModel memoryStrengthsWeaknessesBest for
Apple Silicon workstation
Mac Studio M5 Max / Ultra
96 to 256 GB unifiedLarge models in one address space, silent, low watts, MLX native, no driver workSlower prefill than NVIDIA; no CUDA; vLLM ecosystem limitedSingle operator running 30B to 120B MoE agents, private data
Apple Silicon laptop
MacBook Air/Pro 16 to 48 GB
~10 to 36 GB usablePortable, zero setupOnly small models (3B to 14B, small MoE); swap kills speedAutocomplete, drafts, offline fallback
Consumer NVIDIA rig
1 to 4x RTX 4090/5090
24 to 32 GB per cardFastest tok/s per dollar, CUDA, vLLM/SGLang, fine-tuningVRAM split across cards, power (450 to 600 W each), noise, heatThroughput, batching, LoRA training, serving a team
Pro NVIDIA workstation
RTX PRO 6000 96 GB, DGX Spark class
96 to 128 GBBig models on CUDA in one device, datacenter toolingHighest price per GB; Spark-class boxes have modest bandwidthCUDA-first shops, research, fine-tuning 70B class
AMD unified APU
Ryzen AI Max (Strix Halo) mini PCs
up to ~96 GB of 128 GBCheapest route to large memory, Linux friendlyROCm/Vulkan maturity, lower bandwidth, fewer tuned kernelsBudget large-model experimentation
Rented GPU (hybrid)
Vast.ai, RunPod, H100/H200
80 to 141 GB per GPUNo CapEx, burst to any sizeData leaves premises, hourly cost compounds, cold startsOccasional heavy jobs; never PII

Sizing rule: largest model file + 15 to 20 GB for KV cache and OS = minimum memory tier. A 84 GB model therefore needs the 128 GB tier.

03
Runtimes

MLX / mlx-lm

Apple

Apple's array framework. Fastest path on Apple Silicon, native quantization, OpenAI-style server via mlx_lm.server.

llama.cpp

Universal

C/C++ engine, GGUF format. Runs on Metal, CUDA, ROCm, Vulkan, CPU. The portable fallback and the engine under many GUIs.

Ollama

Convenience

One-command pull and serve, model registry, localhost API on :11434. Trades tuning control for simplicity.

LM Studio

GUI

Desktop app with MLX and llama.cpp backends, model browser, local server. Good for evaluating quants side by side.

vLLM / SGLang

NVIDIA serving

PagedAttention, continuous batching, tensor parallel. The choice when many requests hit one model.

Open WebUI / agents

Front end

Chat UI, RAG and tool calling on top of any OpenAI-compatible endpoint; agent frameworks point here too.

04
Models and sizing

Mixture-of-experts dominates local: large total capacity, 3B to 12B active parameters per token, so speed stays usable.

ModelQuantized sizeRoleMinimum tier
Qwen3-Coder 30B~18 GBFast coder36 GB
Qwen3.6-35B-A3B~33 GBDaily coder64 GB
Nemotron 3.5 Lightning 30B~35 GBAgent driver (Hermes)64 GB
Qwen3.5-122B-A10B~64 GBLong context (262k)96 GB
gpt-oss-120B~65 GBGeneral reasoning96 GB
Nemotron 3 Super 120B~84 GBHeavy reasoning128 GB

Sizes are approximate for 4-bit class quants; long contexts add KV cache on top.

05
E2B: sandboxes for agents

E2B is open-source infrastructure that gives an AI agent a secure, disposable Linux computer. The LLM decides what to run; E2B runs it somewhere that cannot touch your host.

Isolation

Firecracker

Each sandbox is a microVM (the tech behind AWS Lambda). Starts in well under a second; its own kernel boundary.

SDKs

Python / JS

Sandbox.create(), run commands, read/write files, stream stdout, expose ports, kill on exit.

Code Interpreter

Jupyter kernel

Stateful cells, charts returned as data. The data-analysis agent pattern.

Desktop

Computer use

Sandbox with a graphical desktop and stream; agent clicks and types inside the VM, not on your machine.

Templates

Custom images

Bake dependencies into a template so each session starts ready; persistence and pause/resume for long jobs.

Deployment

Cloud or self-host

Managed cloud with usage billing and a free tier; open-source infra can be self-hosted (Terraform, GCP/AWS) when data must stay in-house.

# the canonical loop: model proposes, sandbox executes
from e2b_code_interpreter import Sandbox
from openai import OpenAI                     # pointed at a LOCAL endpoint

llm = OpenAI(base_url="http://localhost:1234/v1", api_key="local")
with Sandbox.create() as sbx:
    code = llm.chat.completions.create(model="qwen3-coder-30b",
            messages=[{"role":"user","content":"Load data.csv and plot revenue by week"}]
           ).choices[0].message.content
    result = sbx.run_code(code)                # runs in the microVM, not on the host
    print(result.logs, result.error)

When to use E2B: any agent that writes and runs code, installs packages, browses, or operates a desktop. When not: pure chat, retrieval, or classification with no execution.

06
Sandbox options compared

OptionBoundaryStartupGPUData stays localFit
E2B cloudFirecracker microVMsub-secondNoNoFastest route to safe agent execution
E2B self-hostedFirecracker microVMsub-secondNoYes (your cloud)Same API, compliance-bound data
Docker on the AI boxContainer (shared kernel)~1 sPossibleYesTrusted code, lowest cost, weaker isolation
gVisor / Kata locallyUser-space kernel / light VM1 to 2 sLimitedYesStronger local isolation, more ops work
Daytona / ModalManaged containers / VMssub-second to secondsModal: yesNoDev environments, GPU functions
No sandboxNone0n/aYesNever for model-written code on a production host

07
Reference stack: local brain, sandboxed hands

Operator / trigger→Agent loop (Hermes)→Local LLM on MLX :1234→Tool call→E2B sandbox (exec)→Result back to agent

Tier A: private

Local model + Docker/gVisor on the same box. Nothing leaves. Use for anything with PII.

Tier B: hybrid

Local model + E2B cloud for execution on non-sensitive code and public data. Best safety per hour of setup.

Tier C: compliance

Local model + self-hosted E2B in own VPC. Strong isolation and residency; highest ops load.

Rule: prompts with client PII never reach a cloud model or cloud sandbox. Route by data class, not by convenience.

08
Verdict

Local AI: Mac Studio M5 Max 128 GB is the floor for a six-model MoE lineup up to 84 GB; MLX first, llama.cpp as fallback, Ollama/LM Studio for convenience. Move to NVIDIA only when throughput or fine-tuning becomes the bottleneck.

E2B: adopt as the execution layer for any agent that runs code. Start on the cloud tier with synthetic or public data; switch to self-hosted or a local gVisor sandbox before any production data enters the loop.

Together: the model thinks locally, the agent acts in a box. Capacity decides the hardware, data class decides the sandbox.

09
Deep dives: three folders

Click a folder to open it; click it again to close.

16 GB unified leaves roughly 10 to 11 GB for weights plus KV cache after macOS. Target models of 8B or less at 4-bit, or small MoE. Memory bandwidth sets speed: M4 ~120 GB/s, M5 ~150 GB/s (M5 also adds GPU neural accelerators that speed up prefill in MLX).

Model (4-bit)SizeM4 Air 16 GBM5 Pro-chassis 16 GBUse
Qwen3 4B / Qwen3.5 4B class~2.5 GB~40 to 55 tok/s~50 to 65 tok/sFast drafts, tool calling
Phi-4-mini 3.8B~2.3 GB~45 tok/s~55 tok/sReasoning, math, short context
Gemma 3 / Gemma 4 small (4B to E4B)~3 GB~35 tok/s~45 tok/sMultilingual ES/EN, vision
Qwen3 8B / Llama 3.1 8B~4.5 to 5 GB~20 to 25 tok/s~25 to 30 tok/sBest quality that fits comfortably
Qwen3-Coder 30B-A3B (3-bit)~13 GBswaps, unusableborderline, close all appsNot recommended at 16 GB

Throughput figures are indicative ranges for MLX 4-bit at short context; verify on your own box with mlx_lm.generate --verbose.

MLX (first choice)

pip install mlx-lm
mlx_lm.server --model mlx-community/Qwen3-8B-4bit --port 1234

Similar runtimes

  • LM Studio: GUI, MLX + GGUF backends
  • Ollama: one-command, llama.cpp under the hood
  • llama.cpp: Metal, GGUF, most quant options
  • Apple Foundation Models: ~3B on-device model via Swift, free, private

16 GB survival rules

  • Raise GPU wired limit: sudo sysctl iogpu.wired_limit_mb=12288 (resets on reboot)
  • Keep context at 8k to 16k; KV cache grows fast
  • Close browsers before loading; watch memory pressure, not free RAM
  • Air throttles on long runs (fanless); Pro chassis sustains

Verdict

Good for autocomplete, private drafts, offline fallback and an agent that calls a bigger remote brain. Not a replacement for the 128 GB box. If buying new: 24 to 32 GB is the real laptop floor.

10
Daily research brief: 2026-10-09

Updated every night by 04:44 Bogota from arXiv, Apple, NVIDIA, Google, Microsoft Research and official release feeds. Last edition .

Local AI Runtimes Advance, Phone AI Gets Multimodal, and Agent Sandboxes Enhance Security

  • Ollama v0.40.2 upgrades models for improved performance and llama.cpp compatibility, while llama.cpp release b11514 includes a Musa FWHT fix and broad platform support. NVIDIA DGX Spark is becoming available with 64GB of unified memory, supporting local AI development, and experiments on a local DGX
  • Google AI Edge Gallery 1.0.20 and LiteRT-LM v0.18.0 both feature official support for EmbeddingGemma 2, enabling on-device multimodal semantic search with NPU and GPU acceleration. The Apache 2.0 license for EmbeddingGemma 2 is noted for its benefits in applications requiring many embedding vectors.
  • E2B sandboxes are used by ClickUp for running AI code on sensitive data with isolated microVMs, and E2B Secrets are introduced for safe credential injection. OpenAI "rogue" agent activities were found on Wikimedia projects, highlighting security challenges. Cowork has shifted its model inference and

Open today's full brief · Permalink · RSS

11
Frequently asked questions

Can a MacBook Air M4 with 16 GB RAM run a local LLM?

Yes, models up to about 8B parameters at 4-bit (roughly 5 GB) run well with MLX or llama.cpp. Qwen3 8B runs around 20 to 25 tokens per second on an M4; 30B-class models swap and are not practical at 16 GB.

What is the best runtime for local AI on Apple Silicon?

MLX (mlx-lm) is usually the fastest on Apple Silicon. LM Studio offers MLX and GGUF backends in a GUI, Ollama is the simplest command-line option, and llama.cpp is the portable fallback.

Mac Studio or NVIDIA DGX Spark for local AI?

Mac Studio M5 Max or Ultra generates tokens faster because of higher memory bandwidth. DGX Spark and GB10 OEM boxes process long prompts faster and run the full CUDA stack, which matters for fine-tuning and CUDA-only code.

What are the DGX Spark OEM alternatives?

ASUS Ascent GX10, Dell Pro Max with GB10, HP ZGX Nano, Lenovo ThinkStation PGX, Acer Veriton GN100 and MSI EdgeXpert use the same GB10 Grace Blackwell chip with 128 GB unified memory; AMD Strix Halo mini PCs are the lower-cost non-NVIDIA alternative.

How much memory do I need to run a 120B model locally?

Take the quantized model file size and add 15 to 20 GB for KV cache and the OS. An 84 GB 120B-class quant needs the 128 GB tier; a 65 GB model fits in 96 GB with short context.

Which phones run local AI models best?

Phones with 12 to 16 GB RAM and strong NPUs: Pixel 10 Pro (Tensor G5 with Google TPU), Galaxy S25/S26 Ultra and other Snapdragon 8 Elite phones, and iPhone 17 Pro. They run 1B to 4B models such as Gemma E2B, Qwen3 4B and Phi-4-mini.

What is Gemma E2B?

Gemma E2B is Google's effective-2B edge variant of Gemma, designed for phones and laptops, with text, image and audio input. It runs offline through Google AI Edge Gallery, LiteRT-LM and MediaPipe. It is unrelated to the E2B sandbox company.

What is E2B and why use it with AI agents?

E2B is open-source infrastructure that gives AI agents secure, disposable Linux sandboxes built on Firecracker microVMs. The agent runs generated code inside the sandbox instead of on your machine, via Python or JavaScript SDKs.

Can E2B be used with a local LLM?

Yes. E2B is model-agnostic: a local model served by MLX, Ollama or LM Studio on an OpenAI-compatible endpoint decides what code to run, and the E2B sandbox executes it. Self-host E2B when data must not leave your infrastructure.

Is local AI cheaper than cloud GPUs?

For daily use it usually is. Renting a 96 GB GPU at about 1 to 2 USD per hour for 4 hours a day breaks even with a 128 GB workstation in roughly 9 to 34 months, and local keeps private data on premises.