Topic 8: Deliverables
10 min read·22 Sept 2026
The README
The README is written for a stranger with ten minutes: install, choose a provider, run over stdio, run over HTTP with auth, run the tests, and fix the usual problems. We followed it step by step in a fresh shell, on port 8112 instead of 8000 to avoid clashing with other services, and it worked as written (the HTTP question returned the SM-2 answer, and the server saw four POSTs). Here is its full text.
text
# notes-assistant
Ask questions about a folder of Markdown research notes. An MCP server exposes the notes
(search, read, create); a small host connects an LLM to that server, asks before any write,
logs every tool call without logging note contents, and warns when the server's tools change.
Tested with Python 3.11 and `mcp==2.2.0` (protocol revision 2026-07-28).
## 1. Install (2 minutes)
```bash
git clone <this repository> notes-assistant && cd notes-assistant
python3.11 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
export PYTHONPATH=. # every command below runs from the repository root
```
## 2. Choose an LLM provider (2 minutes)
Pick one. Groq and Gemini have free tiers; Ollama runs on your machine with no key.
| Provider | Set these | Default model |
|---|---|---|
| Groq (default) | `export GROQ_API_KEY=...` | `llama-3.3-70b-versatile` |
| Gemini | `export LLM_PROVIDER=gemini GEMINI_API_KEY=...` | `gemini-2.5-flash` |
| Ollama | `export LLM_PROVIDER=ollama` and run `ollama pull qwen3:8b` | `qwen3:8b` |
Override the model with `LLM_MODEL`. No key at all? Add `--scripted` to the commands below
to use a deterministic stand-in model (it proves the wiring, not the answer quality).
## 3. Ask a question over stdio (1 minute)
```bash
python examples/m11_cli.py "How many participants will the nap study have?"
```
The CLI launches the server as a subprocess (`python -m notes_assistant.server`), lists its
tools, pins their definitions in `.pins/notes.json`, and runs the agent loop. Tool calls are
logged to stderr; the answer is printed to stdout. If the model wants to create a note, you
see the exact arguments and must type `y`.
Use your own notes with `--notes-dir /path/to/notes` (one `.md` file per note, optional
frontmatter with `title`, `tags`, `created`).
## 4. Run over HTTP with authentication (3 minutes)
Terminal 1, the server:
```bash
export NOTES_AUTH_KEY=change-me-to-a-long-random-string
export NOTES_AUTH_ISSUER=http://127.0.0.1:9000
export NOTES_RESOURCE_URL=http://127.0.0.1:8000/mcp
NOTES_TRANSPORT=streamable-http python -m notes_assistant.server
```
Terminal 2, the host (same three `NOTES_AUTH_*` exports first):
```bash
export NOTES_TOKEN=$(python examples/m11_dev_token.py)
python examples/m11_cli.py --http http://127.0.0.1:8000/mcp "Which algorithm does Anki use?"
```
`m11_dev_token.py` stands in for a real authorization server. The server accepts a token only
if its signature, issuer, expiry, and audience (`NOTES_RESOURCE_URL`) all match, and creating
notes needs the `notes:write` scope (`m11_dev_token.py --write`). In production, point the
verifier at your identity provider and never share the signing key.
## 5. Run the tests (1 minute)
```bash
pytest -q
```
16 tests: the server surface (two tools, the template, the prompt), five answerable questions and one the notes cannot answer (the host must abstain),
a note with an injected instruction (no write without approval, and an approved write is
logged), a token for a different server (rejected over real HTTP), pinning, logs free of note
contents, and stdio, HTTP, and CLI smoke tests. They use a scripted model, so they need no key.
To run the same six questions against your real model:
```bash
python examples/m11_eval_real.py
```
## Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
| `Set GROQ_API_KEY in your environment to use groq.` | No key for the chosen provider | Export the key, switch `LLM_PROVIDER`, or add `--scripted` |
| `ModuleNotFoundError: No module named 'notes_assistant'` | Not running from the repository root | `cd` to the root and `export PYTHONPATH=.` |
| `Could not talk to the server` and `HTTP 401 ... invalid_token` | Token expired (10 minutes), or minted for another URL | Mint a new token with the same `NOTES_RESOURCE_URL` the server uses, including port and `/mcp` |
| `HTTP 401` although the token is fresh | Server and token disagree on issuer or key | Export identical `NOTES_AUTH_*` values in both terminals |
| `WARNING ... pin mismatch server=notes tool=... kind=changed` and the tool disappears | The server's tool definition changed since you pinned it | Read the change (see `docs/WHAT_BROKE.md`); if you trust it, delete `.pins/notes.json` to re-pin |
| `ERROR mcp.client.stdio: Failed to parse JSONRPC message from server`, or the host hangs | Something in the server wrote to stdout | Log to stderr only; never `print()` in a stdio server |
| `Address already in use` | Port 8000 taken | Set `NOTES_PORT` and use the same port in `NOTES_RESOURCE_URL` and `--http` |
| The answer is `I could not find this in your notes.` | Search found nothing relevant, or the model was cautious | Check the question's keywords exist in a note; this is the intended abstain answer |
Layout: `notes_assistant/` (store, server, host, auth, pinning, llm, config), `examples/`
(scripts per course module), `tests/`, `notes/` (sample data), `docs/WHAT_BROKE.md`.Code explained
- In simple words: the one page someone needs to go from "git clone" to a working, tested assistant.
- What happens: section 1 installs the pinned dependencies and sets
PYTHONPATH. Section 2 is a table rather than prose, because choosing a provider is a lookup. Section 3 uses the CLI over stdio and says what the person will see (logs on stderr, ay/Nprompt for writes). Section 4 needs two terminals and sets the same threeNOTES_AUTH_*values in both, which prevents the most common failure we hit (audience and key mismatches). Section 5 runs the tests and names what they cover, and points to the real-model evaluation. The troubleshooting table maps the exact error text a person will see to its cause and fix; every row comes from something that actually happened while building this module. - Comes out: a
README.mdof about 100 lines at the repository root.