A modular, multi-model framework for building AI agents with pluggable skills, persistent memory, and multi-provider orchestration (Gemini · Claude · Ollama).
Most agent setups lock you into a single provider and lose context between runs. This framework abstracts the model layer so the same agent can run on Gemini, Claude, or a local Ollama model, route each step to the best/cheapest provider, and keep persistent memory across sessions.
- 🔌 Multi-model orchestration — a rule-based router sends each step to the best available provider behind one async interface. Reasoning/code → Claude, cheap/bulk → Gemini, local/private → Ollama, with per-step overrides and graceful degradation when a provider's credentials are missing.
- 🧩 Modular skills — drop-in tools the agent can call. Two are built in:
web_fetch(fetch a URL and extract readable text) andfiles(sandboxed read/write/list within a workspace). Each skill is a small, self-contained, independently testable class. - 🧠 Persistent memory — conversations and per-namespace key/value state survive restarts via a zero-dependency SQLite backend (stdlib
sqlite3). An ephemeral in-memory backend is available for tests and one-off runs. Namespaces isolate agents/sessions that share one store. - 💸 Token optimization — budget-aware context trimming keeps the system prompt plus the most recent turns under a configurable token budget, dropping the oldest turns first. Estimation uses a fast
chars/4heuristic, with an exact path via Anthropic'scount_tokens. Anthropic calls cache the system prefix (cache_control: ephemeral) for prompt-caching savings. - 📊 Observability — a built-in
Traceremits one structured record per step to the stdlibloggingsystem and, whenAGENTS_TRACE_FILEis set, to a JSONL trace file (provider, model, latency, token usage, tool calls, errors). No external tracing dependency required.
flowchart LR
U[User / CLI] --> O[Orchestrator]
O --> R{Model Router}
R -->|reasoning / code| C[Claude]
R -->|cheap / bulk| G[Gemini]
R -->|local / private| L[Ollama]
O --> S[Skill Registry]
S --> T1[Skill: web_fetch]
S --> T2[Skill: files]
O --> M[(Persistent Memory:<br/>SQLite / in-memory)]
O --> TR[Tracer / JSONL]
Each step flows through the orchestrator, which trims context to the token budget, routes to an available provider, runs the chat (executing any requested skills in a bounded tool loop), persists the turns to memory, and emits a trace record.
Python 3.11+ · asyncio (fully async core) · provider SDKs (anthropic, google-generativeai) · httpx (Ollama + the web skill) · stdlib sqlite3 for memory · python-dotenv for config · packaged with Hatchling.
git clone https://github.com/AntoniRomera/ai-agents-framework.git
cd ai-agents-framework
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]" # runtime + dev (ruff, pytest)
cp .env.example .env # add your API keys here (never commit .env)
# Run a task (routes to the best available provider for the step kind):
python -m agents.run --task "Summarize the latest news on small language models"
# Or use the installed console script:
agents run --task "What is retrieval-augmented generation?" --statsNo API keys? The Ollama provider needs none — start a local Ollama server and force fully-local routing:
agents run --task "Write a haiku about caching" --kind localagents run --task "<prompt>" [--kind reasoning|cheap|bulk|local|private|code]
[--provider anthropic|gemini|ollama]
[--namespace NAME] [--system "..."]
[--no-tools] [--stats]
agents providers # show which providers are available
agents memory --namespace default # inspect persisted conversation
agents memory --list-namespaces
agents memory --namespace default --clearimport asyncio
from agents import Agent, StepKind
async def main():
agent = Agent("researcher", system="You are concise and cite sources.")
result = await agent.run("Explain prompt caching in one paragraph.", kind=StepKind.REASONING)
print(result.output)
print(result.usage.total_tokens, "tokens")
asyncio.run(main())All configuration is read from the environment (and a .env file if present). See .env.example for the full list.
| Variable | Description | Default |
|---|---|---|
ANTHROPIC_API_KEY |
Claude provider key (provider disabled if unset) | — |
GEMINI_API_KEY |
Gemini provider key (GOOGLE_API_KEY accepted as fallback) |
— |
OLLAMA_HOST |
Local Ollama endpoint | http://localhost:11434 |
OLLAMA_MODEL |
Default Ollama model | llama3.1 |
MEMORY_BACKEND |
sqlite (persistent) or memory (ephemeral) |
sqlite |
AGENTS_DB_PATH |
SQLite database path | ./agents_memory.db |
AGENTS_ROUTER_DEFAULT |
Provider to prefer when a step kind allows fallback | — |
AGENTS_TOKEN_BUDGET |
Max estimated tokens kept per request | 12000 |
AGENTS_TRACE_FILE |
If set, append JSONL step traces here | — |
AGENTS_WORKSPACE |
Sandbox root for the files skill |
current dir |
Model defaults (ANTHROPIC_OPUS_MODEL, ANTHROPIC_SONNET_MODEL, GEMINI_FLASH_MODEL, GEMINI_PRO_MODEL) can also be overridden via the environment.
pip install -e ".[dev]"
ruff check . # lint
ruff format --check . # formatting
pytest -q # test suiteCI runs lint, format check, and the test suite on Python 3.11 and 3.12 (see .github/workflows/ci.yml).
agents/
orchestrator.py # plan + run multi-step tasks (context, routing, tools, memory, tracing)
agent.py # high-level named-agent wrapper
router.py # rule-based provider selection
context.py # token estimation + budget-aware trimming
tracing.py # structured per-step traces (logging + JSONL)
config.py # environment-driven Settings
types.py # shared dataclasses / enums
errors.py # exception hierarchy
cli.py / run.py # command-line entry points
providers/ # anthropic, gemini, ollama (+ Protocol base)
skills/ # registry + web and files skills
memory/ # sqlite and in-memory backends (+ Protocol base)
tests/ # pytest suite (offline; network mocked with respx)
- Streaming responses surfaced through the CLI
- Multi-agent collaboration via
Orchestrator.delegate(sub-agents already supported in the library) - Additional skills (shell, code execution) and a Postgres memory backend
- Example notebook / demo GIF
MIT © Antoni Romera Luis — see LICENSE.
⚠️ Before pushing: ensure no API keys,.envfiles, or private prompts are committed (check the full git history, not just the latest commit).