Part 6: Capabilities
MCP as an interop layer
So far our tools are Python functions in the same process. The Model Context Protocol (MCP) is an open protocol for exposing tools and context to LLM applications, so a tool written once can be used by any MCP-capable host (a chat app, an IDE, your agent). The current specification is version 2026-07-28 at modelcontextprotocol.io/specification/2026-07-28. In its terms:
- A host is the LLM application (our agent). A client is the connector inside the host that talks to one server. A server provides capabilities.
- Messages are JSON-RPC 2.0. Servers offer tools (functions the model can call), resources (data for context), and prompts (templates for users).
- Transports are stdio (a local subprocess) and Streamable HTTP.
The 2026-07-28 revision made the protocol stateless. Per its changelog: the initialize handshake is gone and every request carries its protocol version and client capabilities in _meta; servers must implement a server/discover method; protocol-level sessions (and the Mcp-Session-Id header) are removed, so servers that need state across calls return explicit handles as ordinary tool arguments; every result carries a resultType ("complete", or "input_required" for the new multi round-trip pattern); tasks moved into an optional extension; and Roots, Sampling, and Logging are deprecated. If you read MCP tutorials from 2025, expect them to show the older handshake.
Here is our tool registry answering MCP-shaped requests. It is an in-process teaching shim, not a server: no transport, no authorization, three methods. It exists so you can read real MCP traffic and see how an MCP tool maps onto the tool specs our loop already uses.
"""Module 8: what our tools look like on the wire as MCP (protocol version 2026-07-28).
This is an in-process teaching shim, not a compliant MCP server: it has no
transport, no authorization, and only three methods. It shows the message shapes
so you can read real MCP traffic and see how any MCP server's tools become the
same tool specs our agent loop already uses. For real servers use an official SDK.
"""
from __future__ import annotations
import json
from pathlib import Path
from m08_agent import BillingStore, Tool, make_tools
PROTOCOL_VERSION = "2026-07-28"
READ_ONLY = {"search_kb", "read_article", "get_account", "get_invoice"}
def to_mcp_tool(tool: Tool) -> dict:
"""An MCP Tool definition. Annotations are hints for the host, and hosts must treat them as untrusted."""
return {"name": tool.name, "description": tool.description, "inputSchema": tool.parameters,
"annotations": {"readOnlyHint": tool.name in READ_ONLY,
"destructiveHint": tool.irreversible,
"idempotentHint": tool.name in READ_ONLY or tool.name == "issue_refund",
"openWorldHint": False}}
class McpShim:
"""Answers server/discover, tools/list, and tools/call from our tool registry."""
def __init__(self, tools: dict[str, Tool]) -> None:
self.tools = tools
def handle(self, request: dict) -> dict:
rid, method, params = request["id"], request["method"], request.get("params", {})
if method == "server/discover":
result = {"supportedVersions": [PROTOCOL_VERSION], "capabilities": {"tools": {"listChanged": False}},
"instructions": "Brightlane support tools. issue_refund needs human approval in the host."}
elif method == "tools/list":
result = {"tools": [to_mcp_tool(t) for t in self.tools.values()], "ttlMs": 300000, "cacheScope": "private"}
elif method == "tools/call":
tool = self.tools.get(params.get("name"))
if tool is None: # protocol error: the model cannot fix an unknown tool name by retrying arguments
return {"jsonrpc": "2.0", "id": rid, "error": {"code": -32602, "message": f"Unknown tool: {params.get('name')}"}}
try:
value = tool.fn(**params.get("arguments", {}))
result = {"content": [{"type": "text", "text": json.dumps(value)}], "structuredContent": value, "isError": False}
except Exception as exc: # tool execution error: returned to the model so it can self-correct
result = {"content": [{"type": "text", "text": f"{type(exc).__name__}: {exc}"}], "isError": True}
else:
return {"jsonrpc": "2.0", "id": rid, "error": {"code": -32601, "message": f"Method not found: {method}"}}
return {"jsonrpc": "2.0", "id": rid, "result": {"resultType": "complete", **result}}
def mcp_to_chat_spec(mcp_tool: dict, server_prefix: str) -> dict:
"""Turn an MCP tool into a Chat Completions tool spec, prefixed to avoid clashes between servers."""
return {"type": "function", "function": {"name": f"{server_prefix}__{mcp_tool['name']}",
"description": mcp_tool.get("description", ""),
"parameters": mcp_tool["inputSchema"]}}
if __name__ == "__main__":
Path("runs/m08").mkdir(parents=True, exist_ok=True)
server = McpShim(make_tools(BillingStore("runs/m08/mcp_billing.json"), flaky_failures=0))
meta = {"io.modelcontextprotocol/protocolVersion": PROTOCOL_VERSION,
"io.modelcontextprotocol/clientInfo": {"name": "brightlane-agent", "version": "0.8"}}
print(json.dumps(server.handle({"jsonrpc": "2.0", "id": 1, "method": "server/discover", "params": {"_meta": meta}})))
listed = server.handle({"jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {"_meta": meta}})
for t in listed["result"]["tools"]:
hints = t["annotations"]
print(f" {t['name']:18} readOnly={hints['readOnlyHint']!s:5} destructive={hints['destructiveHint']}")
ok = server.handle({"jsonrpc": "2.0", "id": 3, "method": "tools/call",
"params": {"_meta": meta, "name": "search_kb", "arguments": {"query": "duplicate charge refund"}}})
print(json.dumps(ok)[:220])
bad = server.handle({"jsonrpc": "2.0", "id": 4, "method": "tools/call",
"params": {"_meta": meta, "name": "get_invoice", "arguments": {"invoice_id": "INV-2026-999999"}}})
print(json.dumps(bad))
unknown = server.handle({"jsonrpc": "2.0", "id": 5, "method": "tools/call", "params": {"_meta": meta, "name": "delete_account"}})
print(json.dumps(unknown))
spec = mcp_to_chat_spec(listed["result"]["tools"][0], "brightlane")
print("as a chat tool spec:", spec["function"]["name"])Code explained
- In simple words: dress our seven tools in MCP's clothes, send four requests, and turn one MCP tool back into a chat tool spec.
- What happens:
to_mcp_toolbuilds an MCP tool definition:name,description,inputSchema(our JSON Schema, unchanged), andannotations. The annotation hints (readOnlyHint,destructiveHint,idempotentHint,openWorldHint) are in the spec's schema; the spec also says clients must treat annotations as untrusted unless the server is trusted. That is why our gate keys on our ownirreversibleflag, not on a server's hint.McpShim.handleanswersserver/discover(supported versions and capabilities),tools/list(with thettlMsandcacheScopecaching fields this revision requires on list results), andtools/call.- The spec's two error kinds are both here: a tool execution error (
isError: trueinside a normal result) for problems the model can fix, like a bad invoice id, and a protocol error (a JSON-RPCerrorwith code -32602) for an unknown tool. mcp_to_chat_specis the interop point: any server's tools become Chat Completions tool specs, prefixed with a server name because the spec only guarantees names are unique within one server.
- Comes out:
{"jsonrpc": "2.0", "id": 1, "result": {"resultType": "complete", "supportedVersions": ["2026-07-28"], "capabilities": {"tools": {"listChanged": false}}, "instructions": "Brightlane support tools. issue_refund needs human approval in the host."}}
search_kb readOnly=True destructive=False
read_article readOnly=True destructive=False
get_account readOnly=True destructive=False
get_invoice readOnly=True destructive=False
issue_refund readOnly=False destructive=True
add_internal_note readOnly=False destructive=False
escalate_to_human readOnly=False destructive=False
{"jsonrpc": "2.0", "id": 3, "result": {"resultType": "complete", "content": [{"type": "text", "text": "[{\"article_id\": \"billing-refunds\", \"title\": \"Refunds and cancellations\", \"score\": 4.044}]"}], "structuredCo
{"jsonrpc": "2.0", "id": 4, "result": {"resultType": "complete", "content": [{"type": "text", "text": "ValueError: No invoice 'INV-2026-999999'"}], "isError": true}}
{"jsonrpc": "2.0", "id": 5, "error": {"code": -32602, "message": "Unknown tool: delete_account"}}
as a chat tool spec: brightlane__search_kbThe destructive flag lands only on issue_refund. The failed invoice lookup comes back as a normal result with isError: true, which the host should pass to the model so it can correct itself; the unknown tool is a protocol error.
For real servers, use an official MCP SDK (the Python one is the mcp package on PyPI); check that its version supports the protocol revision your hosts speak, because this revision changed the wire format.
| Situation | Use this | Why |
|---|---|---|
| Tools used only by this one agent, in one codebase | Plain function tools (as in m08_agent.py) | No protocol overhead, easiest to test |
| The same tools should serve several hosts (IDE, chat app, agents) | An MCP server | Write once, connect anywhere |
| Using third-party tools | An MCP client in your host, with your own approval gate and allowlist | You get their tools; you keep control of what runs |
| Stateful tools (a cart, an open browser) | Explicit handles returned by a create tool, passed back as arguments | The 2026-07-28 spec has no protocol sessions |
Code execution as a general-purpose tool
One tool that runs code replaces dozens of narrow tools: proration math, CSV reshaping, date arithmetic, checking a regex. Models are unreliable at arithmetic (Module 1) and good at writing short programs, so "write code, run it, read the output" is often more accurate than "think harder". The cost is risk: you are executing text the model wrote, and that text can be steered by anything in its context, including a malicious ticket (Module 11).
Here is a subprocess runner with time, CPU, memory, and file-size limits, plus a workspace-confined file tool. The last two cases show what these limits do not stop.
"""Module 8: code execution and workspace tools, with limits, and a demonstration of what the limits do NOT stop.
Linux or macOS only (uses the `resource` module). This is a teaching sandbox: it
bounds time, memory, output, and file size. It is NOT an isolation boundary.
"""
from __future__ import annotations
import os
import resource
import subprocess
import sys
import tempfile
import time
from pathlib import Path
def run_python(code: str, timeout_s: float = 2.0, cpu_s: int = 2, mem_mb: int = 256,
file_mb: int = 1, max_output: int = 2000) -> dict:
"""Run untrusted Python in a child process with resource limits. Returns a result dict; never raises."""
def limit_child() -> None: # runs in the child between fork and exec
resource.setrlimit(resource.RLIMIT_CPU, (cpu_s, cpu_s))
resource.setrlimit(resource.RLIMIT_AS, (mem_mb * 2**20, mem_mb * 2**20))
resource.setrlimit(resource.RLIMIT_FSIZE, (file_mb * 2**20, file_mb * 2**20))
os.setsid() # own process group, so a timeout kill takes children with it
with tempfile.TemporaryDirectory(prefix="agent-sbx-") as tmp:
started = time.perf_counter()
try:
proc = subprocess.run([sys.executable, "-I", "-S", "-c", code], cwd=tmp, env={"PATH": "/usr/bin:/bin"},
capture_output=True, text=True, timeout=timeout_s, preexec_fn=limit_child)
outcome = "ok" if proc.returncode == 0 else f"exit {proc.returncode}"
if proc.returncode < 0:
outcome = f"killed by signal {-proc.returncode}"
stdout, stderr = proc.stdout, proc.stderr
except subprocess.TimeoutExpired as exc:
outcome = f"timeout after {timeout_s} s"
stdout = (exc.stdout or b"").decode() if isinstance(exc.stdout, bytes) else (exc.stdout or "")
stderr = ""
ms = round((time.perf_counter() - started) * 1000)
return {"outcome": outcome, "ms": ms, "stdout": stdout[:max_output],
"stderr_tail": stderr.strip().splitlines()[-1][:160] if stderr.strip() else ""}
class Workspace:
"""File tools confined to one directory: the agent's scratch area for drafts and exports."""
def __init__(self, root: str | Path, max_bytes: int = 100_000) -> None:
self.root = Path(root).resolve()
self.root.mkdir(parents=True, exist_ok=True)
self.max_bytes = max_bytes
def _path(self, relative: str) -> Path:
p = (self.root / relative).resolve() # resolve() follows symlinks and removes ..
if not p.is_relative_to(self.root):
raise PermissionError(f"{relative!r} is outside the workspace")
return p
def write_file(self, path: str, content: str) -> dict:
if len(content.encode()) > self.max_bytes:
raise ValueError(f"content over {self.max_bytes} bytes")
p = self._path(path)
p.parent.mkdir(parents=True, exist_ok=True)
p.write_text(content, encoding="utf-8")
return {"written": str(p.relative_to(self.root)), "bytes": len(content.encode())}
def read_file(self, path: str) -> dict:
p = self._path(path)
return {"path": path, "content": p.read_text(encoding="utf-8")[: self.max_bytes]}
def list_files(self) -> list[str]:
return sorted(str(p.relative_to(self.root)) for p in self.root.rglob("*") if p.is_file())
if __name__ == "__main__":
repo_file = str(Path("data/tickets.jsonl").resolve())
cases = {
"proration math": "seats, monthly, annual = 24, 12, 10\n"
"print('monthly per year:', seats * monthly * 12)\n"
"print('annual per year :', seats * annual * 12)\n"
"print('saving :', seats * (monthly - annual) * 12)",
"infinite loop": "while True:\n pass",
"memory bomb": "x = bytearray(1024 * 1024 * 1024)\nprint(len(x))",
"disk filler": "open('big.bin', 'wb').write(b'0' * 5 * 1024 * 1024)",
"reads outside cwd": f"print(open({repo_file!r}).readline()[:60])",
"env and network": "import os, socket\nprint(sorted(os.environ))\n"
"s = socket.create_connection(('pypi.org', 443), timeout=3)\nprint('connected:', s.getpeername()[1])",
}
for name, code in cases.items():
r = run_python(code)
print(f"{name:18} {r['outcome']:22} {r['ms']:>5} ms | out: {r['stdout'].strip().replace(chr(10), '; ')[:84]!r} | err: {r['stderr_tail'][:70]!r}")
ws = Workspace("runs/m08/workspace")
print("\n", ws.write_file("drafts/T-1001.md", "Hi, we refunded the duplicate charge."))
print(ws.list_files())
for bad in ["../../supportdesk/llm.py", "/etc/passwd", "drafts/../../escape.txt"]:
try:
ws.read_file(bad)
except PermissionError as exc:
print("blocked:", exc)Code explained
- In simple words: run model-written Python in a child process with a stopwatch and a memory cap, and give the agent a folder it cannot climb out of.
- What happens:
run_pythonstartspython -I -S -c code(isolated mode, no site-packages) in a fresh temporary directory with an almost empty environment.limit_childruns in the child before it starts:RLIMIT_CPUcaps CPU seconds,RLIMIT_AScaps address space (memory),RLIMIT_FSIZEcaps the size of any file it writes, andsetsidgives it its own process group.subprocess.run(timeout=...)caps wall time. Output is truncated tomax_outputcharacters so a chatty program cannot flood the context.Workspaceresolves every path (following symlinks and..) and refuses anything outside its root. It also caps write size.- The demo runs six programs: a legitimate calculation, an infinite loop, a 1 GiB allocation, a 5 MB file write, a read of a file outside the sandbox directory, and a network connection.
- Comes out: (real runs; timings vary by machine and run)
proration math ok 68 ms | out: 'monthly per year: 3456; annual per year : 2880; saving : 576' | err: ''
infinite loop timeout after 2.0 s 2028 ms | out: '' | err: ''
memory bomb exit 1 49 ms | out: '' | err: 'MemoryError'
disk filler exit 1 67 ms | out: '' | err: 'OSError: [Errno 27] File too large'
reads outside cwd ok 54 ms | out: '{"id": "T-1001", "subject": "Charged twice this month", "bod' | err: ''
env and network ok 107 ms | out: "['LC_CTYPE', 'PATH']; connected: 443" | err: ''
{'written': 'drafts/T-1001.md', 'bytes': 37}
['drafts/T-1001.md']
blocked: '../../supportdesk/llm.py' is outside the workspace
blocked: '/etc/passwd' is outside the workspace
blocked: 'drafts/../../escape.txt' is outside the workspaceThe first four rows are the limits working: the answer comes back in tens of milliseconds, the loop dies at the 2 s timeout, the allocation fails with MemoryError under the 256 MB cap, and the write fails at the 1 MB file limit. The last two rows are the honest part. The child read a ticket file by absolute path, because nothing stopped it; and it opened a TCP connection to pypi.org on port 443, because resource limits do nothing about the network. Stripping environment variables hid the proxy settings but did not remove network access on this machine.
So this is a resource limiter, not a sandbox. It stops accidents (runaway loops, memory blowups). It does not stop a hostile program from reading secrets on disk or sending them somewhere. For code a model writes from untrusted input, run it where there is nothing to steal and nowhere to send it: a container or microVM with no credentials mounted, a read-only filesystem except a scratch directory, and networking off or restricted to an allowlist. Hosted code-execution tools from model providers and dedicated sandbox services exist for exactly this (see Other Tools).
| Situation | Use this | Why |
|---|---|---|
| Model-written code over data you trust, on a dev machine | Subprocess with resource limits (as above) | Stops accidents cheaply |
| Model-written code in production, or any untrusted input in context | Container or microVM, no secrets, no network, disposable | Resource limits do not stop exfiltration |
| A fixed calculation (proration, tax) | A normal function tool | Deterministic, testable, no code generation at all |
| The agent needs files (drafts, exports) | A workspace-confined file tool with size limits | Path traversal is blocked in one place |
Computer and browser use
Computer use means the model operates a graphical interface: it receives a screenshot, returns an action (click at x,y; type text; scroll; press keys), your code performs it, takes a new screenshot, and the loop continues. It is the same perceive, decide, act, observe loop, with pixels as observations. Browser use is the same idea restricted to a web browser, often with page structure (the accessibility tree or DOM) in addition to screenshots.
What is on offer as of September 2026 (check the docs, this moves quickly):
- Anthropic's Claude API documents a computer-use toolset,
computer_toolset_20260801, which its docs describe as generally available with no beta header, with actions such asscreenshot,zoom, clicks,type,key,scroll, andwait. - Google's Gemini API documents computer use on several models, including
gemini-3.5-flash(this course's Gemini default), in a loop where the model returns a function call with coordinates and your code executes it and returns a new screenshot. Responses can include asafety_decisionthat requires you to ask the end user for confirmation. - OpenAI's Responses API documents a
computertool with its own guidance on isolation and confirmation for consequential actions.
When to reach for it: only when there is no API. For Brightlane, the billing system has an API, so a computer-use agent clicking through the billing admin panel would be slower, costlier (every screenshot is image tokens, Module 12), and far less reliable than get_invoice. It earns its place for legacy back-office screens with no API, or for testing your own web UI.
The risks are the ones this module keeps returning to, amplified. Anything on screen is untrusted input: a web page can contain text that looks like instructions. Anthropic's docs recommend a dedicated virtual machine or container with minimal privileges, keeping sensitive data such as logins away from the model, limiting internet access to an allowlist of domains, and asking a human to confirm decisions with real-world consequences such as financial transactions. That is the approval gate from Part 5, applied to clicks.
File and workspace manipulation
Coding agents, report builders, and data agents live in a filesystem: they read inputs, write drafts, run code, and read results. The Workspace class above is the minimum safe version: one root directory, every path resolved and checked, writes capped. Two habits matter more than the code:
- Give the agent its own workspace, never your home directory or the repository root. Our demo blocked
../../supportdesk/llm.py,/etc/passwd, and a traversal hidden inside a relative path. - Treat files the agent reads as untrusted content, like retrieved documents (Module 7). A file can carry injected instructions just as a web page can.
Subagents and delegation
A subagent is an agent started by another agent, with its own fresh context, its own tools, and a narrow task, which returns only a short result. In our code, a subagent is just run_agent called from inside a tool function, with a different system prompt and a subset of tools. Part 7 builds exactly that. Delegation buys three things: a clean context for the subtask (no leftover history competing for attention), a smaller tool list (fewer definitions, easier tool selection, Module 6), and the option to run subtasks in parallel. It costs a hand-off: whatever the subagent needs has to be written into its brief, and whatever it learned has to fit in its report.