Sep 10, 2026
EN PT-BR

SIGNAL FILE 02

AgentForge

STATUS
WORK IN PROGRESS
CATEGORY
OPEN SOURCE

OVERVIEW

AgentForge answers a question that framework documentation cannot: what does the same agent actually look like in each of them? It holds around ninety runnable examples across ten frameworks — AG2, Agno, AutoGen, CrewAI, Google ADK, LangGraph, LlamaIndex, the OpenAI Agents SDK, Pydantic AI, and smolagents — each example isolating one concept.

THE PROBLEM

Choosing an agent framework usually means reading ten marketing pages that all promise the same things. What is missing is the boring comparison: the same task, the same tools, written each way, so the differences are in the code instead of in the pitch.

APPROACH

Every implementation returns through a shared agent contract (`AgentResult` and `TokenUsage`) and calls the same shared tool implementations, which is what keeps the benchmark from quietly favouring whichever framework got the better-written tools. Results are written as JSONL and rendered into Markdown reports. Unit tests run with no credentials and no network, using mocked HTTP, so the repository stays reproducible for anyone who clones it.

FEATURES

  • Capability matrix tracking tools, streaming, memory, human-in-the-loop, multi-agent, RAG, MCP, tracing, and evals across frameworks
  • Benchmark harness that pins the model, temperature, prompts, and tools, then measures latency, token usage, tool calls, and failure rate
  • Each example is a self-contained project with pinned dependencies
  • CI runs linting, compilation, offline unit tests, and secret scanning on every change

STACK

Python / uv / pytest / LLMs / RAG / MCP / Streamlit