ToolsComparison
Top 10 agent orchestration frameworks in 2026, ranked for production
Ten code-first frameworks for single- and multi-agent systems, ranked on orchestration model, durable state, human approval, MCP and A2A support, language reach and licence.

Agent orchestration frameworks are the libraries that decide what happens between model calls: which agent runs next, where state is kept, when a human must approve a step, and how a run resumes after a crash. In October 2026 there are more than a dozen credible code-first options, and they no longer differ much on the basics. Every one in this list calls tools over the Model Context Protocol, streams output and offers some form of approval. What still separates them is the orchestration model, how durable a run is, which languages they serve, and how far they tie you to one model vendor.
This is a ranked list for engineers choosing one for production. It is the companion to our survey of agent frameworks in 2026, which maps the category into camps and explains the protocols; this piece ranks ten specific frameworks against explicit criteria, adds four the survey did not cover (Mastra, LlamaIndex Workflows, Strands Agents and Temporal), and ends with recommendations keyed to constraints. If you would rather not run the orchestration yourself, our sibling ranking of managed agent platforms covers the hosted side.
Every version number and release date below was checked against the project’s GitHub releases or official documentation on 5 October 2026.
How we ranked
We scored each framework on eight criteria, weighted towards what fails in production rather than what demos well.
- Orchestration model. Graph, loop with handoffs, crews of role agents, or durable workflow. We rewarded frameworks that offer an explicit model for multi-step control flow and let you drop to plain code when the model gets in the way.
- State and durability. Whether a run can be checkpointed, resumed after a process restart, and replayed, and at what granularity.
- Human-in-the-loop. Whether approval is a first-class pause that can outlive the process, or a callback that only works while the process is alive.
- MCP and A2A support. MCP client support is now universal; MCP server exposure and Agent2Agent (A2A) client and server support are not.
- Language reach. Python only, or also TypeScript, .NET, Go or Java.
- Model-agnosticism. Whether non-default providers are first-class or a best-effort adapter.
- Production maturity. A stable 1.x API, release cadence, documented deployment and observability.
- Licence. Permissive open source scores highest; split licences and non-standard terms score lower.
We did not run the frameworks against each other. Rankings reflect documented capability and maturity, not measured task accuracy.
The contenders at a glance
| Rank | Framework | Type and licence | Languages | Deployment | Best fit |
|---|---|---|---|---|---|
| 1 | LangGraph | State-graph orchestrator; MIT | Python, TypeScript | Self-hosted, or LangSmith Deployment | Long-running, resumable workflows with human steps |
| 2 | Microsoft Agent Framework | Agents plus graph workflows; MIT | .NET, Python, Go (preview) | Self-hosted, or Microsoft Foundry | .NET and Azure estates; AutoGen and Semantic Kernel migrations |
| 3 | OpenAI Agents SDK | Agent loop with handoffs; MIT | Python, TypeScript | Self-hosted; Temporal for durability | Teams standardised on OpenAI models, voice and sandboxes |
| 4 | Google ADK | Agents plus graph workflows; Apache-2.0 | Python, TypeScript, Go, Java, Kotlin | Agent Runtime, Cloud Run, GKE or containers | Polyglot teams, Google Cloud |
| 5 | Pydantic AI | Typed agent library; MIT | Python | Any; seven-plus durable engines | Python services that want types and engine-backed durability |
| 6 | CrewAI | Crews inside event-driven Flows; MIT | Python | Self-hosted, or CrewAI AMP | Role-decomposed tasks and fast prototypes |
| 7 | Temporal | Durable execution engine; MIT | Go, Java, Python, TypeScript, .NET and more | Self-hosted, or Temporal Cloud | Runs that must survive hours or days of failure |
| 8 | Mastra | TypeScript agents and workflows; Apache-2.0 with ee/ licence | TypeScript | Node apps: Next.js, Express, Astro, SvelteKit | TypeScript product teams |
| 9 | LlamaIndex Workflows | Event-driven workflows; MIT | Python | Self-hosted, FastAPI, llama-agents and llamactl | Retrieval-heavy agents and document pipelines |
| 10 | Strands Agents | Model-driven loop with multi-agent patterns; Apache-2.0 | Python, TypeScript | AWS services or any container | AWS-centred teams that want a small loop |
The top 10, ranked
1. LangGraph
LangGraph describes itself as “a low-level orchestration framework and runtime for building, managing, and deploying long-running, stateful agents”, and that is an accurate summary of why it ranks first. You build a StateGraph of nodes that read and write shared state, connect them with edges (including conditional ones), and attach a checkpointer so every step is persisted to a thread. The LangGraph overview lists durable execution, human-in-the-loop, short- and long-term memory, and streaming as its four core capabilities, and names Klarna, Uber and J.P. Morgan among users.
Strengths. Human approval is built on the persistence layer, not beside it. A node calls interrupt(), LangGraph saves the graph state and waits, and Command(resume=...) continues it with the human’s answer as the interrupt’s return value. Because the pause is a saved checkpoint, it can last as long as your database does. The deployment server is now an A2A peer: the Agent Server A2A endpoint at /a2a/{assistant_id} implements “A2A v1.0 JSON-RPC binding” while still accepting v0.3 method names, publishes an Agent Card, and maps A2A contextId to a LangGraph thread_id. The same server exposes graphs as MCP tools at /mcp. The Python package reached 1.2.12 on 21 September 2026, and the repository is MIT licensed.
Limitations. The graph is ceremony when the task is a simple loop. The interrupt semantics catch people out: “The node restarts from the beginning of the node where the interrupt was called when resumed, so any code before the interrupt runs again”, so side effects before an interrupt must be idempotent. And the most polished deployment path, LangSmith Deployment, is a commercial product; self-hosting the open-source server is possible but you own Postgres, Redis and upgrades.
Best fit: workflows with branches, long pauses and human sign-off, where you want state you can inspect and resume from any step.
2. Microsoft Agent Framework
Microsoft Agent Framework is the merger of AutoGen and Semantic Kernel. The overview calls it “the direct successor, created by the same teams”, combining “AutoGen’s simple agent abstractions with Semantic Kernel’s enterprise features” and adding graph workflows. The 1.0 announcement on 3 April 2026 declared agents, middleware, memory, workflows and multi-agent orchestration stable “with backward compatibility commitments”. Releases since then are frequent: python-1.20.0 on 2 October and dotnet-1.23.0 on 1 October 2026.
Strengths. It has the widest catalogue of named multi-agent patterns of any framework here: Sequential, Concurrent, Handoff, Group Chat and Magentic, the last being “A manager agent dynamically coordinates specialized agents”. Every orchestration can pause for a human through approval-required tools. Workflows checkpoint at superstep boundaries so long runs survive interruption. First-party connectors cover Microsoft Foundry, Azure OpenAI, OpenAI, Anthropic Claude, Amazon Bedrock, Google Gemini and Ollama. A2A is documented in .NET, Python and Go, including an A2AExecutor that exposes any agent as an A2A server. It is the only serious option for .NET teams.
Limitations. The 1.0 post said “A2A 1.0 support coming soon”, and the Python A2A package still installs with --pre. The Go port is in public preview, and “Declarative agents, RAG, CodeAct, and functional workflows are not yet available” there. DevUI, Foundry hosting and the Claude Code SDK integration were all labelled preview at 1.0. The documentation’s quickstarts assume Foundry, and its legal note puts responsibility for data flowing to non-Azure models on you.
Best fit: .NET or Azure organisations, and anyone migrating an AutoGen or Semantic Kernel codebase.
3. OpenAI Agents SDK
The OpenAI Agents SDK keeps the surface deliberately small: agents, handoffs and agents-as-tools, guardrails, sessions and tracing, plus sandbox agents that run “inside real isolated workspaces” and realtime voice agents. It describes itself as “Python-first”, using language features rather than new abstractions to chain agents. Python reached v0.23.1 on 2 October 2026; the TypeScript SDK is at v0.18.0. Both are MIT licensed.
Strengths. Human approval is clean and persistent. Mark a tool needs_approval=True, read pending calls from RunResult.interruptions, convert to a RunState, call approve() or reject(), and resume with Runner.run(agent, state); to_json() and from_json() let a pause span process restarts (human-in-the-loop docs). Session backends include SQLite, Redis, SQLAlchemy, MongoDB, Dapr and an encryption wrapper. For crash durability, Temporal’s generally available integration runs each model call as a Temporal Activity “so they retry durably and are not repeated during Workflow replay”, and proxies MCP calls the same way.
Limitations. The version number is still below 1.0. Non-OpenAI models work through ModelProvider or an OpenAI-compatible client, but the LiteLLM and any-llm adapters are offered on a “best-effort, beta basis” (models docs), and tracing needs reconfiguring without an OpenAI key. Hosted tools such as web search and code interpreter tie the agent to OpenAI’s platform. There is no graph; complex branching lives in your code. Under Temporal, LocalShellTool, ComputerTool and SQLiteSession are unsupported.
Best fit: teams standardised on OpenAI models who want the shortest path to handoffs, approvals, voice and sandboxed code execution.
4. Google ADK
Google’s Agent Development Kit is “the open-source agent development framework that lets you build, debug, and deploy reliable AI agents at enterprise scale” (adk.dev). Version 2.0 added graph workflows that “compose both AI-powered agents and deterministic execution nodes into a flexible execution graph that can include decision branching”. The Python package reached v2.11.0 on 2 October 2026 and the TypeScript package 2.2.0 on 30 September, both Apache-2.0.
Strengths. Language reach is unmatched: Python, TypeScript, Go, Java and Kotlin. ADK is also the most complete on protocols, documenting MCP tools, agents exposed as MCP servers, and A2A both for exposing and consuming agents. Human input in graph workflows is a node that yields RequestInput; a rerunOnResume flag decides whether the interrupted node re-executes or passes the reply straight to its successor, which is a more explicit answer to LangGraph’s idempotency trap (human input docs). Evaluation is part of the toolkit, with user simulation, environment simulation and custom metrics. Model support extends beyond Gemini to Claude, LiteLLM, Ollama and vLLM, with runtime model routing.
Limitations. The best-integrated deployment target is Google Cloud’s Agent Runtime; Cloud Run and GKE are documented, but other clouds are a container you run yourself. Gemini is the default path and the most tested. Graph workflows arrived in 2.0 and the older workflow-agent templates remain, so examples in the wild mix two styles. The TypeScript port trails Python in version and community size.
Best fit: polyglot organisations, Java and Kotlin shops, and teams deploying on Google Cloud.
5. Pydantic AI
Pydantic AI is a typed agent library from the team behind Pydantic. An agent is generic over a dependency type and an output type, which the project says moves “whole classes of errors from runtime to write-time”. It supports “virtually every model and provider (OpenAI, Anthropic, Google, Bedrock, Azure AI Foundry, Groq, Mistral, xAI, Ollama, and dozens more), swappable with a string” (overview). It is MIT licensed and ships fast: v2.54.0 on 3 October 2026.
Strengths. Durability is delegated to engines that already do it well. The durable execution page lists five co-maintained integrations (Temporal, DBOS, Prefect, Restate and AWS Lambda) plus Kitaru, Apache Airflow and Absurd, and “Each engine wraps every model request and tool call as its own durable unit”. The documentation draws a distinction other frameworks blur: “Durability is not storage.” Approval uses deferred tools: requires_approval=True ends the run with DeferredToolRequests, and the caller resumes with DeferredToolResults, including argument overrides. Instrumentation is OpenTelemetry-native. It is both an MCP client and server, and A2A exposure goes through FastA2A.
Limitations. Python only. Multi-agent orchestration is plain Python plus an optional pydantic-graph package; there are no named patterns like Agent Framework’s. FastA2A was extracted from Pydantic AI and now lives in the datalayer/fasta2a repository, and Pydantic AI 2 replaced agent.to_a2a() with an agent_to_a2a function, so older examples break. Fast minor releases mean frequent upgrades.
Best fit: Python services that want typed outputs, provider freedom and durability from Temporal or DBOS rather than a framework-specific checkpoint format.
6. CrewAI
CrewAI calls itself “the leading open-source framework for orchestrating autonomous AI agents and building complex workflows”, combining “collaborative intelligence of Crews with the precise control of Flows”. It has an MIT licence and a release every one to two weeks (1.15.23 on 28 September 2026).
Strengths. The role abstraction is the quickest way to prototype a decomposable task: give agents roles, goals and tools, and let a crew delegate. The documentation now recommends wrapping crews in Flows, which add @start, @listen and @router decorators, typed state with a UUID per run, @persist for state across restarts, and @human_feedback, which pauses a flow and can have an LLM collapse free-text feedback into a named outcome. A2A support is first class: an agent can act as an A2A client, delegating to remote agents with timeouts, turn limits and OAuth2 or API-key auth, and as an A2A server over JSON-RPC, gRPC or HTTP+JSON.
Limitations. Crews delegate control to the model, which makes runs less predictable and more expensive in tokens than an explicit graph. The MCP adapter “primarily supports adapting MCP tools”; prompts and resources are not integrated. Persistence defaults to SQLite, which is a poor fit for multi-instance deployments without extra work. Python only. The managed AMP platform is where much of the operational tooling lives.
Best fit: research, content and analysis tasks that split naturally into roles, and teams that want a working multi-agent prototype in an afternoon.
7. Temporal
Temporal is not an agent framework. It is a durable execution engine, MIT licensed (server v1.32.0, Python SDK 1.34.0 on 30 September 2026), and it ranks here because a growing share of production agents run on top of it. Its Durable AI documentation promises that “a Workflow resumes automatically after a crash, a network timeout, or a multi-day wait”, and lists integrations with the Vercel AI SDK, Deep Agents, Google ADK, LangGraph, Mastra, the OpenAI Agents SDK, Pydantic AI, Spring AI and Strands Agents.
Strengths. It solves the hardest production problem, not losing work, better than any framework-native checkpoint. Model and tool calls become Activities with retry policies, workflow history is replayed exactly, and human approval is a Signal the workflow blocks on, for minutes or weeks. SDKs cover Go, Java, Python, TypeScript and .NET, among others. Because it sits under frameworks rather than replacing them, adopting it does not force a rewrite of agent logic. Temporal Cloud removes the operational burden for teams that do not want to run the cluster.
Limitations. The determinism rule is real: workflow code cannot do I/O directly, so every model call, tool call and clock read has to be an Activity, and some framework features (local shell tools, SQLite sessions) do not survive the move. Running a Temporal cluster is a platform commitment. It offers no agent abstractions of its own; you still need one of the frameworks above for prompting, tools and handoffs.
Best fit: organisations that already run Temporal, and agents whose runs last hours or days and must never repeat a side effect.
8. Mastra
Mastra is “a TypeScript framework for building AI agents and applications”, with agents, Zod-typed tools, workflows, memory and evaluation in one package. It releases several times a week; @mastra/core 1.74.0 shipped on 5 October 2026. The licence file puts most of the code under Apache 2.0 and anything in an ee/ directory under a separate enterprise licence.
Strengths. Models are addressed as provider/model strings, so switching providers is a string change. Workflows compose steps with .then() and control-flow methods, share typed workflow state, and can suspend and later resume with resumeStream(); Mastra Studio shows a graph view and “time travel” replay of individual steps. Tool calls can require approval before execute runs, and network approval snapshots execution state so a multi-agent run can pause for a human. MCPClient consumes MCP servers and MCPServer exposes Mastra agents, tools and workflows to other clients. It slots into Next.js, Astro, Express and SvelteKit.
Limitations. TypeScript only. The API moves quickly: agent networks via .network() are already marked deprecated in favour of supervisor agents, and approval snapshots fail without a configured storage provider. The ee/ split means some features are not under the open-source licence. We did not find first-party A2A documentation.
Best fit: TypeScript product teams building agents into a web application.
9. LlamaIndex Workflows
LlamaIndex Workflows is an event-driven orchestration library: “A step receives an event, does some work, and returns another event,” and the returned event triggers whichever step’s type annotation accepts it (docs). The package, llama-index-workflows 2.25.0 (25 September 2026), now lives in the MIT-licensed run-llama/llama-agents repository alongside deployment tooling, and LlamaIndex’s AgentWorkflow and FunctionAgent are built on it.
Strengths. Branching and looping are plain Python, not a DAG definition, which keeps complex retrieval pipelines readable. Durability is well specified: state is “the events still in flight plus the state store”, Context.to_dict() and Context.from_dict() restore a run in a different process, completed steps do not re-run, and in-progress steps are rewound for at-least-once recovery; a DBOS runtime plugin can journal step transitions automatically. Human input is a pair of events, InputRequiredEvent and HumanResponseEvent, and the context can be serialised while waiting. It pairs naturally with LlamaIndex’s retrieval and document-parsing stack.
Limitations. It is Python only in the documentation we read. wait_for_event() “replays all code preceding it”, the same idempotency trap as LangGraph. Multi-agent patterns are thinner than Agent Framework’s or Strands’. We did not find first-party A2A documentation, and deployment tooling (llama-agents, llamactl) is younger than the library.
Best fit: agents whose hard part is retrieval, parsing and document workflows rather than agent-to-agent coordination.
10. Strands Agents
Strands Agents is an open-source SDK from AWS that takes “a model-driven approach to building AI agents in just a few lines of code”. Its repository, now the strands-agents/harness-sdk monorepo, holds Python and TypeScript packages under Apache 2.0; Python reached 1.57.2 on 1 October 2026. The README’s own pitch is precise: “Choose Strands when you would otherwise write your own agent loop: it runs in your process with no hosted control plane.”
Strengths. The core is a small loop with lifecycle controls (turn limits, token budgets, cancellation), but the Strands site documents a full set of multi-agent patterns on top: graph, swarm, workflow, agents as tools, and A2A. MCP servers are first class, and built-in tools cover shell, file and web access. Model support spans Anthropic, OpenAI, Amazon Bedrock, Gemini, Ollama and others. Deployment guides cover Lambda, Fargate, App Runner, EKS and EC2 as well as plain Docker and Kubernetes, and Temporal lists Strands among its integrations for durable runs.
Limitations. Durability is not native; state survives a restart through sessions or an external engine, not checkpointed graph state. The documentation and examples lean towards Bedrock and AWS services. The repository restructuring means older import paths and links in community posts may be stale. It has a smaller community than the frameworks above it.
Best fit: AWS-centred teams that would otherwise hand-roll an agent loop and want multi-agent patterns available when they need them.
Also considered
Four projects came close. AutoGen is in maintenance: its latest tagged release is python-v0.7.5 from 30 September 2025, and Microsoft points users to Agent Framework. AG2, the community fork of AutoGen, is active (v1.1.2 on 3 October 2026, Apache-2.0) and is listed by the A2A project as a supporting framework, but it is smaller than the ten above. Dify is popular (1.17.1 shipped on 10 September 2026) but is a low-code platform rather than a code-first framework, and GitHub reports a non-standard licence. The Claude Agent SDK and smolagents are covered in our agent frameworks survey; the first is tied to Claude models, the second has not tagged a release since May 2026.
What actually differs
Feature lists converge; four design choices do not.
Orchestration model
Graphs (LangGraph, Agent Framework workflows, ADK 2.0, Strands’ graph pattern) make control flow explicit and inspectable. Loops with handoffs (OpenAI Agents SDK, Pydantic AI, Strands’ default) let the model plan and keep the code small. Crews (CrewAI, Agent Framework’s Group Chat and Magentic) delegate coordination to role agents. Durable workflows (Temporal, Mastra workflows, LlamaIndex events with DBOS) treat each call as a recorded step. Most frameworks now offer two of these; the default is what shapes your code.
The evidence favours starting simple. “The Illusion of Multi-Agent Advantage” (Jwalapuram et al., June 2026) found that “automatic MAS consistently underperform CoT-SC despite being up to 10x more expensive”, while expert-designed multi-agent systems did win on tasks built to need decomposition. The practical reading is a single agent first, and a graph or crew only when a pause, a branch or genuine parallelism appears in the requirements.

Figure 1: Four orchestration models cover the top ten. Most frameworks offer more than one, but each has a default shape that decides how your code reads.
Where state lives
| Framework | Unit of state | Resume granularity | Approval survives a restart? |
|---|---|---|---|
| LangGraph | Checkpoint per step on a thread | Node (code before interrupt() re-runs) | Yes |
| Microsoft Agent Framework | Checkpoint per superstep | Superstep | Yes |
| OpenAI Agents SDK | Session plus serialisable RunState | Paused tool call; activity under Temporal | Yes, via to_json() |
| Google ADK | Session and graph workflow state | Node, with rerunOnResume choice | Yes |
| Pydantic AI | Engine-dependent | Model request or tool call | Yes, via deferred tools or engine |
| CrewAI | Flow state with UUID | Flow method | Yes, with @persist |
| Temporal | Workflow event history | Exact Activity | Yes, Signals |
| Mastra | Workflow snapshot | Suspended step | Yes, with storage configured |
| LlamaIndex Workflows | Events in flight plus store | Step, at-least-once | Yes, via Context.to_dict() |
| Strands Agents | Session | Turn | Through sessions or an engine |
The split that matters is between frameworks that checkpoint their own state, which is convenient but proprietary in format, and those that hand calls to an engine, which is more work but gives exactly-once semantics for model calls and a format that outlives the framework.
Protocols
MCP client support is universal in this list. A2A is close behind: LangGraph’s Agent Server, Agent Framework, ADK, CrewAI, Pydantic AI and Strands all document it, and the A2A project, which joined the Linux Foundation’s Agentic AI Foundation on 27 August 2026 at v1.0, names LangGraph, CrewAI, Pydantic AI and AG2 as supporting frameworks. The practical consequence is that tool servers and remote agents are now the portable assets; orchestration code is not.
Language and vendor gravity
Python teams can choose any of the ten except Mastra. TypeScript teams have LangGraph, the OpenAI Agents SDK, ADK, Mastra and Strands. .NET teams have Agent Framework and Temporal. Java and Kotlin teams have ADK and Temporal. Each lab SDK works best with its own models and cloud, and each framework vendor’s managed platform is where operational lock-in now lives.
The infrastructure layer under any framework
Every framework above decides which model to call and which tool to invoke. None of them is designed to decide which provider serves the call when the first one is returning 5xx errors, whether this agent has budget left this month, or whether a given MCP tool should be visible to it at all. Those are infrastructure concerns, and agents make them sharper: a loop that retries a failing call, or a crew that fans out to five tool servers, multiplies cost and blast radius in a way a chat application does not.
The common pattern is to point the framework’s model client and its MCP client at a gateway rather than directly at providers and tool servers. Because every framework in this list accepts an OpenAI-compatible base URL or a custom model provider, and every one speaks MCP, this is a configuration change, not a rewrite. We covered the mechanics in what an MCP gateway does and compared gateways in five LLM gateways compared.
Bifrost, an open-source Go gateway from Maxim AI, is one example that covers both halves. As an LLM gateway it exposes an OpenAI-compatible API over 20+ providers, and fallbacks move a request to the next provider/model entry after the primary exhausts its retries, with each fallback treated as a new request so caching, governance and logging run again. Its published benchmarks report 11 µs of added overhead at 5,000 requests per second on a t3.xlarge instance, which is small next to a model call.
As an MCP gateway, Bifrost acts as an MCP client to upstream tool servers and optionally as an MCP server to agents. By default tool calls are “suggestions only”, and auto-execution has to be opted into per tool; when it is used purely as a gateway, “Bifrost has no LLM loop” and approval stays with the client framework (MCP overview). Virtual keys give each agent or team its own budget, token and request limits, allowed models and allowed MCP tools, which is what makes AI governance enforceable per agent rather than per application. Logs, Prometheus metrics and OpenTelemetry traces cover model and tool calls together, so gateway-level observability does not depend on which framework’s tracing you chose. Bifrost Edge extends the same policies to AI apps and MCP servers on employee machines.

Figure 2: The framework decides what to call; the gateway decides how, and its policy survives a change of framework.
Two points are worth stating. Enterprise features such as clustering and in-VPC deployment come with the paid tier. And a gateway does not replace the framework’s own approval logic; it narrows what an agent can reach, it does not decide whether a human should sign off. For a wider view of the gateway options, Maxim AI’s own comparison of the top five LLM gateways in 2026 covers latency, failover and governance across five products.
Recommendations keyed to constraints
Start from the constraint that is hardest to change.
- Runs wait hours or days for humans or external systems. LangGraph with a Postgres checkpointer is the default. If the organisation already runs Temporal, put the agent on Temporal, using Pydantic AI or the OpenAI Agents SDK integration, and get exactly-once model calls as a bonus.
- You run .NET or live in Azure. Microsoft Agent Framework. Check the A2A and Go packages’ preview status before committing to them.
- You are a TypeScript product team. Mastra if you want agents, workflows and a studio in one package; the OpenAI Agents SDK if you are on OpenAI models; ADK or LangGraph’s TypeScript port if you need graphs.
- You are committed to one model vendor. Use that vendor’s SDK: the OpenAI Agents SDK, ADK for Gemini, or Strands for Bedrock. Accept the portability cost knowingly.
- You must stay provider-neutral in Python. Pydantic AI for typed services, LangGraph for graphs with human steps.
- The task splits naturally into roles. CrewAI, inside a Flow, with a token budget per run.
- The hard part is retrieval and documents. LlamaIndex Workflows.
- Polyglot teams or Java and Kotlin. ADK.
Whichever you pick, route model and MCP traffic through a gateway from day one, so that failover, budgets and tool permissions are configuration rather than framework code.

Figure 3: Four questions narrow ten frameworks to one or two for most teams.
The bottom line
LangGraph leads because it answers the production questions, where state lives, how a run resumes and how a human approves, with the least hand-waving, and it now speaks A2A v1.0. Agent Framework and ADK are close behind for teams whose language or cloud already points at Microsoft or Google. The more important decision may be the one beneath the framework: tool servers on MCP, remote agents on A2A, and a gateway in front of models and tools are the parts that survive the next migration. Treat the framework as replaceable and invest in those.
Feature claims are drawn from public documentation and vendor announcements as of October 2026; check current docs before deciding.
Sources
- LangGraph overview LangChain
- LangGraph interrupts LangChain
- A2A endpoint in Agent Server LangChain
- Microsoft Agent Framework overview Microsoft Learn
- Microsoft Agent Framework version 1.0 Microsoft
- Workflow orchestrations in Agent Framework Microsoft Learn
- Host agents with A2A Microsoft Learn
- OpenAI Agents SDK documentation OpenAI
- OpenAI Agents SDK: models OpenAI
- OpenAI Agents SDK: human in the loop OpenAI
- OpenAI Agents SDK: sessions OpenAI
- Agent Development Kit documentation Google
- ADK: models Google
- ADK graph workflows: human input Google
- Pydantic AI overview Pydantic
- Pydantic AI: durable execution Pydantic
- Pydantic AI: deferred tools Pydantic
- FastA2A repository Datalayer
- CrewAI Flows CrewAI
- CrewAI A2A agent delegation CrewAI
- CrewAI MCP overview CrewAI
- Durable AI Temporal
- OpenAI Agents SDK integration Temporal
- Mastra workflows overview Mastra
- Mastra network approval Mastra
- Mastra licence Mastra
- LlamaIndex Workflows LlamaIndex
- LlamaIndex Workflows: durable workflows LlamaIndex
- LlamaIndex Workflows: human in the loop LlamaIndex
- Strands Agents Strands Agents
- Strands Agents SDK repository Strands Agents
- A new chapter for A2A: joining the Agentic AI Foundation A2A Project
- The Illusion of Multi-Agent Advantage arXiv
- Bifrost MCP gateway overview Maxim AI
- Bifrost fallbacks Maxim AI
- Bifrost virtual keys Maxim AI
- Bifrost benchmarks Maxim AI
Frequently asked questions
What is the best agent orchestration framework in 2026?
For most production teams, LangGraph. It combines an explicit state graph, checkpointed persistence, interrupt-based human approval and an A2A endpoint in its deployment server, under an MIT licence. Microsoft Agent Framework is the better pick for .NET and Azure estates, and Mastra for TypeScript product teams.
Is AutoGen still maintained?
Microsoft describes Microsoft Agent Framework as the direct successor to AutoGen and Semantic Kernel, built by the same teams, and Agent Framework reached 1.0 in April 2026. The microsoft/autogen repository's latest tagged release is python-v0.7.5 from 30 September 2025. AG2 is a separate community fork that still ships releases.
What is the difference between LangGraph and CrewAI?
LangGraph models an agent system as a graph of nodes over shared, checkpointed state, so the developer draws the control flow. CrewAI models it as crews of role-playing agents inside event-driven Flows, so more of the coordination is delegated to the agents. LangGraph gives finer control over resumption; CrewAI gets a role-based prototype running faster.
Do I need Temporal if my framework already has checkpoints?
Not usually. Framework checkpoints cover most approval pauses and restarts. Temporal earns its place when a run must survive hours or days of infrastructure failure, when model and tool calls must never repeat on recovery, or when the organisation already runs Temporal for other workflows.
Which agent frameworks support both MCP and A2A?
As of October 2026, LangGraph (through its Agent Server), Microsoft Agent Framework, Google ADK, CrewAI, Pydantic AI (through FastA2A) and Strands Agents document both. The OpenAI Agents SDK, Mastra and LlamaIndex Workflows document MCP; we did not find first-party A2A documentation for them in the pages we read.
Should I use a single agent or a multi-agent system?
Start with a single agent. A June 2026 arXiv study found automatically designed multi-agent systems underperformed a single model with chain-of-thought and self-consistency while costing up to ten times more. Add agents when the task genuinely decomposes, needs separated context, or benefits from parallel work.
Why put an AI gateway under an agent framework?
Agents multiply model and tool calls, so a provider outage, a runaway loop or an over-permissioned tool hurts more than in a chat app. A gateway applies failover, per-agent budgets, MCP tool allow-lists and logging at one point, independent of which framework issued the call.


