About Portfolio Cases Services Blog Contact 🎙 Talk to AI
EN DE RU
🎙 Talk to AI
October 3, 2026 · 3 min read

AI Agents for Long-Horizon Tasks: Running a Full Dev Team on 5GB VRAM with late-cli

I'm Denis Shokhirev, Agentic AI Systems Architect based in Freiburg. At DennisCraft AI Studio I ship autonomous multi-agent systems to production for DACH B2B clients—my stack: Claude, Supabase, n8n, Doppler, and self-hosted Postgres. On a recent industrial automation deployment, I ran into a hard constraint: ship a production-grade AI dev pipeline, but the only GPU available had 5GB VRAM. This isn't a demo. It's a real-world bottleneck that eats your delivery timeline. The Challenge: Long-Hor

Denis Shokhirev
Denis Shokhirev
Agentic AI Systems Architect
Telegram LinkedIn

I'm Denis Shokhirev, Agentic AI Systems Architect based in Freiburg. At DennisCraft AI Studio I ship autonomous multi-agent systems to production for DACH B2B clients—my stack: Claude, Supabase, n8n, Doppler, and self-hosted Postgres. On a recent industrial automation deployment, I ran into a hard constraint: ship a production-grade AI dev pipeline, but the only GPU available had 5GB VRAM. This isn't a demo. It's a real-world bottleneck that eats your delivery timeline.

The Challenge: Long-Horizon Workloads on Tight GPU Budgets

Most multi-agent AI demos run on cloud A100s or unlimited VRAM. In production, real DACH clients often have only entry-level NVIDIA cards (e.g., T400, T1000, V100)—4 to 5GB VRAM for the entire agent crew. If your agents need to handle 30+ minute codegen or refactoring tasks, standard approaches collapse: you hit out-of-memory, stalling, or deadlocks.

Symptoms: Agent Hangs and Orchestration Stalls

In practice, I've seen agents silently hang on large tasks, n8n flows deadlock, and codegen LLMs (Claude, GPT-4) crash on long files. Generating or reviewing even a 2,000-line Python file is enough to freeze the pipeline if you don't optimize memory management.

Solution Landscape—and Why late-cli Actually Works

Most solutions are either managed LLM APIs with strict token limits (Anthropic, OpenAI) or self-hosted models that need >10GB VRAM. For 5GB VRAM, only a handful of tools work: late-cli (https://github.com/late-labs/late), quantized models, and aggressive pipeline optimization. No magic, no vendor lock-in.

Architecture: Assembling an AI Dev Team on 5GB VRAM

My real stack:

  • Claude Code (API only; no local model)
  • late-cli for agent orchestration and task scheduling
  • Supabase/Postgres—persistent storage for context, tasks, artifacts
  • n8n—pipeline automation (triggers, CI/CD, notifications)
  • Doppler—environment variables and secrets management

late-cli: Memory Management and Task Control

late-cli is an open-source CLI for running LLM agents with minimal RAM/VRAM footprint. Its strength is dynamic VRAM allocation per agent/task, with strict queuing and automatic memory cleanup after each run. You don't need to keep every agent alive simultaneously.

# Launch 3 agents with VRAM cap
late agent run --model=phi-2 --max-vram=4800 --tasks=tasks.yaml

# tasks.yaml:
# - name: "review_code"
#   input: "src/app.py"
# - name: "write_tests"
#   input: "src/app.py"
# - name: "refactor"
#   input: "src/utils.py"

Serializing Tasks, Not Threads

The trick is to schedule agents sequentially, not keep them all in memory. I chunk long tasks (batch size 256–512 tokens). For each, the agent spins up, runs, then releases VRAM. The task queue lives in Supabase, and n8n monitors pipeline state—if an agent fails or stalls, I get notified and can restart just that agent.

Pipeline: How It Works in Production

Step-by-step:

  1. Tasks (issues) are stored in Supabase: context, code, requirements.
  2. n8n triggers agent launch via late-cli, capping VRAM at 4.8GB.
  3. The agent processes a batch and writes output back to Supabase.
  4. n8n checks status, triggers the next batch or agent as needed.
  5. Claude Code API is used for review and critical code suggestions.
  6. Post-completion: CI/CD pipeline, deployment, and security audit (semgrep, bandit).
import requests

def trigger_agent(task_id, input_path):
    response = requests.post(
        "http://localhost:8000/run",
        json={
            "task_id": task_id,
            "input": open(input_path).read(),
            "max_vram": 4800
        }
    )
    return response.json()

Comparison: How Leading LLM Frameworks Behave

FrameworkMin VRAMDynamic Memory ReleaseBatch Control
late-cli4GBYesYes
LangChain8GB+NoLimited
OpenLLM10GB+NoPartial

Security: Catching Vulnerabilities in Agent-Generated Code

A 2024 Stanford CodeML paper (https://arxiv.org/abs/2402.00957) found 38% of LLM-generated Python contained CWE-89 patterns. I've personally caught SQL injection and XSS issues on three recent agent releases—even post-review by Claude.

  • semgrep—for static analysis of Python/TypeScript
  • bandit—for Python vulnerability checks
  • gitleaks—for scanning secrets in repos
semgrep --config=auto src/
bandit -r src/
gitleaks detect --source=./repo

FAQ

Can I run larger models on 5GB VRAM?

Depends on quantization. phi-2 (INT4) runs fine; Llama-3 doesn't, even quantized.

How do you persist long context between agents?

Supabase or Postgres—context is chunked in a separate table, task ID as key.

late-cli vs LangChain for low-resource setups?

late-cli actually frees memory, while LangChain holds more state and eats VRAM.

Debugging: what if an agent hangs?

n8n signals + logging in Supabase. If no response in 10 minutes, auto-kill and restart.

What’s the minimal stack for this?

late-cli, Supabase/Postgres, n8n. Claude Code API is for review—optional.

Where in your production LLM pipeline do agents stall most—data annotation, codegen, or CI/CD? I'd genuinely like to know. I run a free 30-min stack audit for DACH founders building AI in regulated markets. DM me on LinkedIn or write to @ger_dennis_ai.

Continue reading
How to Instantly Spot Technical Debt and Architecture Issues in TypeScript/JS: Automated Codebase Audits with fallow
OpenAI Dots: Your AI Agent Works 24/7 Even When You're Offline — How to Deploy and What Are the Business Risks
AI Agents Now Jailbreak Each Other: Real-World Self-Replicating Prompt Injection and How to Defend Production
Why Your AI Agents Go Dumb or Rogue in Production: Real Fails of Self-Learning and Evolution Loops
All articles →
Where this is applied
Services — what we build
Talk to the voice agent
Case studies
Ready to build?

Turn your process into an AI system

Production quality. DACH B2B focus.

Start a project → ← All articles