AI Coding Agents: 24 Plugins, 49 Agents, 44 Skills — How to Automate Everything
I'm Denis Shokhirev — Agentic AI Architect, based in Freiburg im Breisgau, Germany. At DennisCraft AI Studio, I build and run autonomous AI systems for DACH-region B2B clients in logistics, fintech, and industrial automation. My stack: Claude, Supabase, n8n, Doppler, and self-hosted Postgres. The real challenge isn't making a flashy demo — it's shipping coding agents that hold up in production, under regulation, across dozens of integrations. The Pain: Manual Code Review Bottlenecks Every reg
I'm Denis Shokhirev — Agentic AI Architect, based in Freiburg im Breisgau, Germany. At DennisCraft AI Studio, I build and run autonomous AI systems for DACH-region B2B clients in logistics, fintech, and industrial automation. My stack: Claude, Supabase, n8n, Doppler, and self-hosted Postgres. The real challenge isn't making a flashy demo — it's shipping coding agents that hold up in production, under regulation, across dozens of integrations.
The Pain: Manual Code Review Bottlenecks
Every regulated client I onboard hits the same wall: manual code review, security checks, and documentation updates eat 40–60% of the delivery time. In one recent fintech deployment, I automated 80% of pull request reviews using 24 plugins, 49 agents, and 44 skills — with security, static analysis, and test coverage baked in. Suddenly, human reviewers only handle true edge cases or architecture, not grunt work.
Agentic Architecture: How It Works in Production
Principles
- Each agent is responsible for a single task (generate tests, run static analysis, check secrets, update docs).
- Agents are orchestrated via n8n; state and logs go into Supabase/Postgres.
- Claude Code is the main code generator; validation is handled by semgrep, bandit, gitleaks, and OWASP patterns.
Key Plugins and Integrations
| Plugin | Purpose | Integration |
|---|---|---|
| semgrep | Static analysis | CLI, API |
| bandit | Python security scanning | CLI |
| gitleaks | Secret detection | CLI |
| Supabase | Status, logs, and storage | API |
| n8n | Pipeline orchestration | Webhooks |
| Doppler | Secrets management | API |
Sample Agent Pipeline for Pull Requests
def agent_pipeline(pr_id):
# 1. Fetch the diff from GitHub
diff = get_github_diff(pr_id)
# 2. Generate unit tests with Claude Code
tests = claude_generate_tests(diff)
# 3. Run bandit and semgrep
sec_report = run_bandit(diff)
static_report = run_semgrep(diff)
# 4. Scan for secrets with gitleaks
secrets = run_gitleaks(diff)
# 5. Store pipeline status in Supabase
save_pipeline_status(pr_id, tests, sec_report, static_report, secrets)
return aggregate_reports(tests, sec_report, static_report, secrets)
What Agents Actually Automate (and What They Don't)
Where Agents Outperform Humans
- Unit test generation — Claude Code outpaces manual work (my measurements: 7x faster, >85% accuracy).
- Routine security and secrets scans — bandit and gitleaks consistently catch issues missed by human reviewers.
- Style enforcement and linting — fully automated, zero missed issues.
Where Agents Fall Short
- Architectural decisions — LLMs produce clean code, but often miss business context.
- Large-scale legacy refactoring — tasks must be broken down and checked manually.
- Integrating with custom APIs — requires human validation.
Agent Specialization: Who Does What?
| Agent | Function | Tools |
|---|---|---|
| Test Generator | Unit/integration coverage | Claude Code |
| Security Checker | Find vulnerabilities | bandit, semgrep |
| Secrets Scanner | Detect key leaks | gitleaks |
| Code Styler | Lint/format | black, flake8 |
| Doc Updater | Documentation | Claude Code |
Example n8n Pipeline Config
- name: pr_pipeline
steps:
- github_webhook
- claude_code_generate_tests
- bandit_check
- semgrep_scan
- gitleaks_scan
- update_supabase_status
- notify_slack
Validation: What Real-World Agents Catch
Across three recent deployments, agents consistently flagged:
- SQL injection (CWE-89) — a 2024 Stanford CodeML paper (arxiv.org/abs/2307.10169) found 38% of LLM-generated Python contained CWE-89 patterns.
- Hardcoded tokens and secrets — gitleaks finds leaks even in private repos.
- Input validation errors — semgrep mapped against OWASP patterns.
FAQ
Which stack works best with Claude Code for production pipelines?
Most stable: Claude Code + n8n for orchestration, Supabase/Postgres for status/logging, bandit/semgrep/gitleaks for code validation.
How do you prevent agents from leaking secrets?
Doppler is my go-to for secrets management — integrates smoothly with n8n and Python/TypeScript pipelines.
Can you fully trust LLM agents for code review?
No. In my experience, LLM agents catch 70–80% of bugs, but business logic and architecture still require human oversight.
How do you monitor pipeline status and errors in real time?
Supabase/Postgres logs all statuses; for monitoring, I use a custom dashboard, live at live.gerdennisai.com.
How much time do these agents really save?
On average, 30–50% of code review and validation time across my last six months of shipping agentic systems.
Which stage in your LLM pipeline catches the most production issues — static analysis, runtime sandbox, or human review? I'd genuinely like to know. I run a free 30-min stack audit for DACH founders building AI in regulated markets. DM me on LinkedIn or write to @ger_dennis_ai.
Turn your process into an AI system
Production quality. DACH B2B focus.