RagLeap Packages
41 tests passing in CI on Python 3.10, 3.11 and 3.12 — scripted model only, no real model has been run through this loop yet

ragleap-agents

A small act-observe agent loop over ragleap-tools Tool objects. The model proposes one action at a time and you bring it as a plain callable. No provider code, no database.

0.1.0
PyPI version (alpha)
41
Tests passing
3
Python versions in CI
1
Runtime dependency (ragleap-tools)

Install

pip install ragleap-agents
# or, with uv
uv add ragleap-agents

Quickstart

The model returns one JSON action per step: {"tool": "<name or done>", "arguments": {...}}.

from ragleap_tools import CALCULATOR_TOOL
from ragleap_agents import Agent, Policy, TRUSTED

def my_llm(prompt: str) -> str:
    ...  # call any model, return its text

agent = Agent(llm=my_llm, tools=[CALCULATOR_TOOL],
              policy=Policy(max_steps=4, tools={"calculator": TRUSTED}))
result = agent.run("What is 6 * 7?")
print(result.status, result.answer)

Approval gates: pause and resume

from ragleap_agents import ToolPolicy
policy = Policy(tools={"write_file": ToolPolicy(taints=False, outbound=True, requires_approval=True)})
agent = Agent(my_llm, tools, policy, store=my_store)
result = agent.run("save the report")
if result.status == "awaiting_approval":
    call = result.pending   # {"call_id", "tool", "arguments", "forced"}
    result = agent.resume(result.run_id, {call["call_id"]: True})   # False rejects

A paused run is a plain JSON dict in a StateStore. InMemoryStateStore is included. resume() raises ResumeError for an unknown run, a run not awaiting approval, or a decision that does not match the pending call, so an approval cannot be replayed.

The safety model

  • Every proposal is checked. The tool must exist and the arguments must fit its JSON Schema (required keys, unknown keys, type and enum only). Anything else stops the run with invalid_plan; no tool runs.
  • Tool results are untrusted. Capped at 1,500 characters each (4,000 in the prompt), fenced in <observation> tags, and lookalike tags are neutralised case-insensitively.
  • Taint rule. After a tool declared taints=True has run, every later outbound=True tool needs approval. The policy can only tighten.
  • Fail closed. A tool with no declared policy counts as both tainting and outbound. Use TRUSTED only for tools that neither return external content nor act outside.
  • Limits. max_steps default 4, hard cap 8; duplicate-action stop; a wall-clock deadline checked between calls. A tool's exception text never reaches the model or the result, only its type.

What this does not do

It does not stop prompt injection. It limits what an injected instruction can do, and only as far as your ToolPolicy declarations are accurate. Tool descriptions (capped at 300 characters) are shown to the model, so tools from an untrusted source can inject through them.

Verified

  • 41 tests with a scripted model, no network: passing on Python 3.10, 3.11 and 3.12 in CI.
  • Mutation-checked: removing the taint rule, the case-insensitive fence, the exception-text rule, the resume claim, the hard cap, the fail-closed default or the bool-is-not-integer check each makes at least one test fail.
  • The published 0.1.0 wheel installs in fresh environments on 3.10, 3.11 and 3.12 and passes a scripted smoke test, with file digests matching the publish log.

Not verified

  • Any real model: no provider has been run through this loop yet.
  • Long runs, and concurrent resumes across processes (a store must make resume atomic).
  • Equivalence with RagLeap Core's agent loop, which it mirrors as of 2026-10-03.
  • Any comparison with other agent frameworks: no speed or capability claim is made.