Metadata-Version: 2.4
Name: p53-courier
Version: 0.1.5
Summary: CLI-based AI assisted web querying with configurable breadth and depth
Keywords: search,web,ai,ollama,local-llm,point53,research
Author-email: "Point 53, LLC" <dev@point53.ai>
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-Expression: MPL-2.0
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: End Users/Desktop
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Text Processing :: Markup :: Markdown
License-File: CONTRIBUTING.md
License-File: LICENSE
License-File: NOTICE
License-File: THIRD_PARTY_LICENSES.md
License-File: TRADEMARKS.md
Requires-Dist: click==8.4.2
Requires-Dist: dateparser==1.3.0
Requires-Dist: feedparser==6.0.12
Requires-Dist: httpx==0.28.1
Requires-Dist: mcp==1.28.1
Requires-Dist: ollama==0.6.1
Requires-Dist: p53-collector==0.1.4
Requires-Dist: platformdirs==4.10.0
Requires-Dist: protobuf==7.34.0
Requires-Dist: pydantic==2.13.4
Requires-Dist: selenium==4.41.0
Requires-Dist: tomli==2.4.1; python_version < '3.11'
Requires-Dist: webdriver-manager==4.0.2
Requires-Dist: anthropic==0.116.0 ; extra == "all"
Requires-Dist: anthropic==0.116.0 ; extra == "anthropic"
Project-URL: Homepage, https://point53.ai/courier.html
Provides-Extra: all
Provides-Extra: anthropic

# Point 53 Courier™

**Your question. Any URL-formatted query location. A cited answer.**

Point 53 Courier is a CLI web-querying agent. You give it a query, a *breadth* (links per layer) and a *depth* (layers to follow); Courier navigates the specified web location, scrapes and summarizes each source with local AI, and returns a cited transcript alongside an LLM-written report. It runs local-first: your queries and the harvested content stay on your machine unless you pick a cloud model. It is part of the Point 53 tools.

- **Local-first AI**: curation, summarization, and analysis run on your hardware via [Ollama](https://ollama.com); a cloud-usage warning fires whenever content would leave your machine.
- **Breadth and depth you control**: `--breadth` sets links per layer; `--depth` sets how many layers deep to follow. Start shallow, go deeper when a topic warrants it.
- **Cited, greppable output**: every run is a directory of Markdown + JSONL: a human `report.md`, a per-source `transcript.md`, and a machine-readable `sources.jsonl`.
- **Targeted lens**: `--target` focuses curation, summarization, and analysis toward exactly what you're after.
- **Agent-ready**: ships an MCP server and a Skill, so agents can run research through the same engine you do.

> **License:** [Mozilla Public License 2.0](https://mozilla.org/MPL/2.0/). See `LICENSE`, `NOTICE`, and `THIRD_PARTY_LICENSES.md`.

## Demo

> **WARNING:** The demo contains some instances of fast jump cuts and light changes. Do not watch if you have photosensitive epilepsy.
>
> **NOTE:** Ctrl + click (Cmd + click on macOS) to open the video in a new tab.

[![Point 53 Courier™ Public Alpha Demo: Dispatch, Scout, Deliver](https://img.youtube.com/vi/bIeuf4DRoEI/maxresdefault.jpg)](https://youtu.be/bIeuf4DRoEI)

## Installation

Courier installs as a standalone CLI tool with [uv](https://docs.astral.sh/uv/) from the Point 53 PEP 503 index. (PyPI hosts fail-loud stub packages only, to prevent name squatting; install from `dist.point53.ai`, not PyPI.) It builds on **Point 53 Collector**'s scraping-and-summarization engine, which is pulled in automatically; Collector is a required dependency and can be used separately.

### Linux / macOS

```bash
# Install uv if you don't have it
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install Point 53 Courier
uv tool install p53-courier --index https://dist.point53.ai/simple/
```

### Windows

```powershell
# Install uv if you don't have it
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

# Install Point 53 Courier
uv tool install p53-courier --index https://dist.point53.ai/simple/
```

### Optional extras

Cloud (Anthropic) inference is an opt-in extra; the base install is local-only by design:

```bash
uv tool install 'p53-courier[anthropic]' --index https://dist.point53.ai/simple/
```

## First Run

**Requirements**

- **Python 3.10+**: uv manages this for you.
- **Ollama** running locally or on your network, with the default models pulled:
  ```bash
  ollama pull gemma4:e4b         # summary + post-analysis
  ollama pull qwen3.5:9b         # chat
  ollama pull granite4.1:8b      # search-result links → RSS
  ```
- **Firefox**: the default scraping browser. Selenium drives a separate instance to keep scraping out of your everyday browser's profile, and Courier assumes that everyday browser is Chrome. Adjust the [configuration](#configuration) if Firefox is your default, so Chrome runs for Courier instead.

**Get going**

`uv tool install` does not create config files. After install, create defaults and check the install:

```bash
courier config doctor
courier doctor

# Run your first query (default breadth 5, depth 1)
courier query "open-source LLMs"
```

`courier config doctor` writes `courier.toml` and `models.toml` under `~/.config/point53/courier/` with sensible defaults. `courier doctor` checks config, storage, model reachability, the Collector dependency, and the webdriver. Once both are green you're ready to research.

A later suite host, **Point 53 Handler**, will orchestrate multi-tool workflows. Each tool stays independently installable.

## Quick Start

```bash
# Focused query with a target lens
courier query "open-source LLMs" --target "new open-weight model releases and benchmarks"

# Go deeper and wider, then ask a secondary question of the findings
courier query "AI agent frameworks" --target "compare capabilities" \
  --analysis-prompt "which is best for a small team under 50 people?" --breadth 10 --depth 2

# Different breadth after the first layer
courier query "local inference" --target "setup guides" --breadth 8 --breadth-n 3 --depth 3

# Interactive chat with the report as context
courier query "running models on a laptop" --target "step by step" --chat

# Read and analyze a single URL
courier url "https://docs.vllm.ai/en/latest/" --target "serving and quantization"

# Fully headless for cron / automation; print to stdout
courier query "this week in open models" --full-headless --stdout-only | tee research.txt

# Rebuild the derived index from run output
courier reindex
```

`--depth 2` is a reasonable default; deeper runs cost more time and more tokens.

### Commands

```
courier query "<query>" [flags]   research a query through the configured search engine
courier url <url> [flags]         extract and analyze links from a single URL
courier profile init|status|reset manage authenticated browser sessions (see Authenticated sessions)
courier doctor                    health checks (config, storage, model, Collector, webdriver)
courier config validate           validate config against the Pydantic schema
courier config doctor             validate + create missing config files with defaults
courier reindex                   rebuild courier.db from run output
courier serve-mcp                 stdio MCP server for agents
```

### Flags (query / url)

| Flag | Purpose |
| --- | --- |
| `--target` | Focus lens for curation, summarization, and analysis (recommended with `url`) |
| `--analysis-prompt` | Secondary analysis question asked of the finished transcript + summary |
| `--breadth` | Links per depth layer (default from config) |
| `--breadth-n` | Override breadth for layers after the first |
| `--depth` | Number of depth layers (default from config) |
| `--profile` | Browser profile mode: `sandboxed` (ephemeral, default) or `credentialed` (see [Authenticated sessions](#authenticated-sessions)) |
| `--pause` | Pause on each scraped page for manual interaction — bare `--pause` waits 15 s, `--pause 30` a custom duration; any key continues, `e` extends 10× |
| `--chat` / `--chat-prompt` | Open an interactive chat with the report as context (`--chat-prompt` seeds the first question) |
| `--headless` | Run the summarization browser headless (query / navigation stays visible) |
| `--full-headless` | Run every browser headless (may trip bot detection on search pages) |
| `--stdout-only` | Print to stdout; skip file output and the index |
| `--i-understand-the-risks` | Suppress cloud-model usage warnings for the run |

> `--pause` drives a visible browser, so it's incompatible with `--full-headless`. A standing default lives under `[defaults]` in `courier.toml`.

## Configuration

Two TOML files under `~/.config/point53/courier/`. Run `courier config doctor` to create them with defaults:

- **`courier.toml`**. Operational settings: the `[search]` source (set manually, see below), breadth/depth defaults, browser/webdriver options, and the `[collector.*]` passthrough that overrides Collector's scraping behavior for Courier's runs.
- **`models.toml`**: per-role LLM config (`provider`, `model`, `base_url`, optional `api_key`) for curation, summarization, and chat. Each role independently selects `ollama`, `anthropic`, or `openai-compatible`.

### Search source

The `query` subcommand drives whatever you point it at. There's no built-in search engine; you set `[search].search_url` by hand, and `{query}` is the placeholder Courier URL-encodes your query into at run time. It can be **any** network location whose results are a links page with the query in the URL, exactly what you see in your browser's address bar after searching. Copy that URL, swap your search term for `{query}`, done.

```toml
# ~/.config/point53/courier/courier.toml
[search]
search_url = "http://localhost:8888/search?q={query}"   # default: local SearXNG

# Any endpoint works. Paste your address-bar URL and replace the term with {query}:
# search_url = "https://search.example.com/search?q={query}"   # a public search engine
# search_url = "https://your-intranet.local/search?q={query}"  # any internal/private endpoint
```

The default is a local [SearXNG](https://docs.searxng.org/) instance. Prefer SearXNG when you can: queries stay on your network, you avoid commercial search-result pages that mix organic links with malvertising, and you are not depending on a commercial search engine's terms of service for automated result harvest. Excessive requests to a WAN endpoint may occasionally be rate-limited.

API keys resolve **env-first** (file-last): `P53_<PROVIDER>_API_KEY` > vendor env (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) > the literal `api_key` in `models.toml`. Routing any role to a non-local `base_url` prints a cloud-usage warning; suppress with `--i-understand-the-risks` or in config.

## Authenticated sessions

Some sources sit behind a login. `courier profile` manages a persistent **credentialed** browser profile so Courier can research pages you're signed into, kept separate from the default ephemeral **sandboxed** session.

```bash
# Open a visible browser on the credentialed profile; log in, then press Enter to save
courier profile init                   # optional: --browser chrome | firefox

# Show the credentialed profile's path, size, and last-modified time
courier profile status

# Delete the credentialed profile and every saved session (prompts to confirm)
courier profile reset
```

Log into your sites with **"Remember me"** enabled during `profile init`, then point runs at that profile:

```bash
courier query "internal roadmap discussions" --profile credentialed
```

…or make it the default in `courier.toml`:

```toml
[profile]
default = "credentialed"               # default is "sandboxed" (ephemeral)
```

> **Agents never get your logins.** Credentialed profiles work only through the `courier` CLI; the MCP server is sandboxed-only, so an agent can't drive your authenticated sessions.

## How it works

```
query ─▶ search engine ─▶ LLM curates top --breadth links into RSS
                                  │
                                  ▼
                    scrape + summarize each source (local AI)
                                  │
             depth > 1? ─yes─▶ follow links, curate next layer ─┐
                                  │                              │
                                  └──────────────◀──────────────┘
                                  ▼
       transcript.md + sources.jsonl ─▶ holistic LLM report ─▶ report.md
```

### Risk reduction

Courier (via Collector's shared fetch policy and its own print path) reduces several classes of risk while harvesting the open web:

- **http/https only:** `check_fetch_url` refuses non-http(s) schemes before navigation, and re-vets the final URL after redirects.
- **Private-network notice:** loopback, RFC 1918, and `.local` hosts are allowed for intranet research; Courier prints a notice when a private host is fetched so the exposure is visible.
- **Control-character blocking:** `_strip_ctl` strips C0/C1 controls from untrusted model and web text before it is printed to the terminal, blocking OSC/ANSI injection.

## Output

Each run is a self-describing directory under `~/.local/share/point53/courier/runs/<timestamp>_<slug>/`:

| File | Contents |
| --- | --- |
| `report.md` | Holistic LLM summary, plus an optional `## Post-Analysis` section |
| `transcript.md` | Per-source summaries (from Collector's distill) |
| `sources.jsonl` | One JSON object per source, the machine-readable record |
| `feed-1.xml`, `feed-2.xml`, … | Curated RSS for each depth layer |
| `manifest.json` | Typed manifest (`type: "courier-run"`, schema v1.1, plus enterprise-coordination fields) |

The Markdown is the source of truth; `courier.db` is a derived index, rebuildable any time with `courier reindex`.

## File locations

Default paths by platform. On **macOS** and **Windows**, config and data resolve to the *same* directory.

**macOS with `XDG_*` set:** if `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, or `XDG_CACHE_HOME` is exported (common in dotfile-managed setups), Point 53 follows it instead of the `~/Library` default for that column, independently per variable. With the conventional values this collapses to the Linux/XDG layout, e.g. config at `~/.config/point53/courier/`.

| Root | Linux | macOS | Windows |
| --- | --- | --- | --- |
| Config | `~/.config/point53/courier/` | `~/Library/Application Support/point53/courier/` | `%LOCALAPPDATA%\point53\courier\` |
| Data | `~/.local/share/point53/courier/` | `~/Library/Application Support/point53/courier/` | `%LOCALAPPDATA%\point53\courier\` |
| Cache | `~/.cache/point53/courier/` | `~/Library/Caches/point53/courier/` | `%LOCALAPPDATA%\point53\Cache\courier\` |

Config holds `courier.toml` and `models.toml`; data holds `courier.db` and the `runs/` output; cache holds the ephemeral `storage.db` and scratch briefings (wiped after each run).

Operational logs are shared across the suite under the state root: `~/.local/state/point53/logs/` (`~/Library/Application Support/point53/logs/` on macOS, `%LOCALAPPDATA%\point53\logs\` on Windows).

Paths shown elsewhere in this README use the Linux/XDG form.

## Agent surface

- **MCP**: `courier serve-mcp` exposes `query` and `url` as stdio MCP tools, with schemas auto-derived from the same Pydantic models as the CLI.
- **Skill**: `src/p53_courier/skills/courier/SKILL.md` for agent runtimes without MCP.

## TODO

**Planned**

- [ ] OS-keyring-backed API keys (beta-2): drop the plaintext key option from `models.toml`
- [ ] Token / cost accounting for cloud model roles
- [ ] Fix `courier doctor` not detecting installed web browsers on macOS and Windows
- [ ] Improved workflow and clarity when chaining Point 53 Courier runs
- [ ] Decision-model AI for next-stage site interaction paths: a controlled, browser-based alternative to computer-use processes (beta-2)
- [ ] Remove `--keep-open` (and `[defaults].keep_open`) and make the post-run revisit menu the default for interactive runs — slated as the test version bump after the inaugural signed release

**Known limitations (alpha)**

- Deeper `--depth` / wider `--breadth` runs cost more time and tokens; there's no hard budget cap yet.
- `--full-headless` can trip bot detection on some pages.
- Hard dependency on Point 53 Collector's engine (installed automatically): Courier can't run without it.

## License and Attribution

- **Point 53 Courier source code:** Mozilla Public License, v. 2.0. See `LICENSE`.
- **Third-party runtime dependencies and external tools:** see `NOTICE` for a summary and `THIRD_PARTY_LICENSES.md` for full attribution.
- **Ollama, Anthropic, and other provider models:** not distributed with Courier; each carries its own license from its respective publisher.
- **Trademarks:** "Point 53" and "Point 53 Courier" are trademarks of Point 53, LLC. MPL-2.0 section 2.3 excludes trademark rights from the copyright/patent grant; nothing in the license authorizes use of these marks. The CLI command `courier` is a functional identifier, not a trademark claim on the generic English word.

