AI Agents for Long-Horizon Tasks: Running a Full Dev Team on 5GB VRAM with late-cli
I'm Denis Shokhirev, Agentic AI Systems Architect based in Freiburg. At DennisCraft AI Studio I ship autonomous multi-agent systems to production for DACH B2B clients—my stack: Claude, Supabase, n8n, Doppler, and self-hosted Postgres. On a recent industrial automation deployment, I ran into a hard constraint: ship a production-grade AI dev pipeline, but the only GPU available had 5GB VRAM. This isn't a demo. It's a real-world bottleneck that eats your delivery timeline. The Challenge: Long-Hor
I'm Denis Shokhirev, Agentic AI Systems Architect based in Freiburg. At DennisCraft AI Studio I ship autonomous multi-agent systems to production for DACH B2B clients—my stack: Claude, Supabase, n8n, Doppler, and self-hosted Postgres. On a recent industrial automation deployment, I ran into a hard constraint: ship a production-grade AI dev pipeline, but the only GPU available had 5GB VRAM. This isn't a demo. It's a real-world bottleneck that eats your delivery timeline.
The Challenge: Long-Horizon Workloads on Tight GPU Budgets
Most multi-agent AI demos run on cloud A100s or unlimited VRAM. In production, real DACH clients often have only entry-level NVIDIA cards (e.g., T400, T1000, V100)—4 to 5GB VRAM for the entire agent crew. If your agents need to handle 30+ minute codegen or refactoring tasks, standard approaches collapse: you hit out-of-memory, stalling, or deadlocks.
Symptoms: Agent Hangs and Orchestration Stalls
In practice, I've seen agents silently hang on large tasks, n8n flows deadlock, and codegen LLMs (Claude, GPT-4) crash on long files. Generating or reviewing even a 2,000-line Python file is enough to freeze the pipeline if you don't optimize memory management.
Solution Landscape—and Why late-cli Actually Works
Most solutions are either managed LLM APIs with strict token limits (Anthropic, OpenAI) or self-hosted models that need >10GB VRAM. For 5GB VRAM, only a handful of tools work: late-cli (https://github.com/late-labs/late), quantized models, and aggressive pipeline optimization. No magic, no vendor lock-in.
Architecture: Assembling an AI Dev Team on 5GB VRAM
My real stack:
- Claude Code (API only; no local model)
- late-cli for agent orchestration and task scheduling
- Supabase/Postgres—persistent storage for context, tasks, artifacts
- n8n—pipeline automation (triggers, CI/CD, notifications)
- Doppler—environment variables and secrets management
late-cli: Memory Management and Task Control
late-cli is an open-source CLI for running LLM agents with minimal RAM/VRAM footprint. Its strength is dynamic VRAM allocation per agent/task, with strict queuing and automatic memory cleanup after each run. You don't need to keep every agent alive simultaneously.
# Launch 3 agents with VRAM cap
late agent run --model=phi-2 --max-vram=4800 --tasks=tasks.yaml
# tasks.yaml:
# - name: "review_code"
# input: "src/app.py"
# - name: "write_tests"
# input: "src/app.py"
# - name: "refactor"
# input: "src/utils.py"
Serializing Tasks, Not Threads
The trick is to schedule agents sequentially, not keep them all in memory. I chunk long tasks (batch size 256–512 tokens). For each, the agent spins up, runs, then releases VRAM. The task queue lives in Supabase, and n8n monitors pipeline state—if an agent fails or stalls, I get notified and can restart just that agent.
Pipeline: How It Works in Production
Step-by-step:
- Tasks (issues) are stored in Supabase: context, code, requirements.
- n8n triggers agent launch via late-cli, capping VRAM at 4.8GB.
- The agent processes a batch and writes output back to Supabase.
- n8n checks status, triggers the next batch or agent as needed.
- Claude Code API is used for review and critical code suggestions.
- Post-completion: CI/CD pipeline, deployment, and security audit (semgrep, bandit).
import requests
def trigger_agent(task_id, input_path):
response = requests.post(
"http://localhost:8000/run",
json={
"task_id": task_id,
"input": open(input_path).read(),
"max_vram": 4800
}
)
return response.json()
Comparison: How Leading LLM Frameworks Behave
| Framework | Min VRAM | Dynamic Memory Release | Batch Control |
|---|---|---|---|
| late-cli | 4GB | Yes | Yes |
| LangChain | 8GB+ | No | Limited |
| OpenLLM | 10GB+ | No | Partial |
Security: Catching Vulnerabilities in Agent-Generated Code
A 2024 Stanford CodeML paper (https://arxiv.org/abs/2402.00957) found 38% of LLM-generated Python contained CWE-89 patterns. I've personally caught SQL injection and XSS issues on three recent agent releases—even post-review by Claude.
- semgrep—for static analysis of Python/TypeScript
- bandit—for Python vulnerability checks
- gitleaks—for scanning secrets in repos
semgrep --config=auto src/
bandit -r src/
gitleaks detect --source=./repo
FAQ
Can I run larger models on 5GB VRAM?
Depends on quantization. phi-2 (INT4) runs fine; Llama-3 doesn't, even quantized.
How do you persist long context between agents?
Supabase or Postgres—context is chunked in a separate table, task ID as key.
late-cli vs LangChain for low-resource setups?
late-cli actually frees memory, while LangChain holds more state and eats VRAM.
Debugging: what if an agent hangs?
n8n signals + logging in Supabase. If no response in 10 minutes, auto-kill and restart.
What’s the minimal stack for this?
late-cli, Supabase/Postgres, n8n. Claude Code API is for review—optional.
Where in your production LLM pipeline do agents stall most—data annotation, codegen, or CI/CD? I'd genuinely like to know. I run a free 30-min stack audit for DACH founders building AI in regulated markets. DM me on LinkedIn or write to @ger_dennis_ai.
Turn your process into an AI system
Production quality. DACH B2B focus.