Metadata-Version: 2.4
Name: p53-intercept
Version: 0.2.6
Summary: Point 53 Intercept — ad-hoc evaluation of verbal and visual data with offline transcription and local LLM analysis.
Keywords: audio,transcription,whisper,faster-whisper,ollama,local-llm,point53
Author-email: "Point 53, LLC" <dev@point53.ai>
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-Expression: MPL-2.0
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: End Users/Desktop
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Multimedia :: Sound/Audio :: Capture/Recording
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
License-File: CONTRIBUTING.md
License-File: LICENSE
License-File: NOTICE
License-File: THIRD_PARTY_LICENSES.md
License-File: TRADEMARKS.md
Requires-Dist: click==8.4.2
Requires-Dist: httpx==0.28.1
Requires-Dist: mcp==1.28.1
Requires-Dist: ollama==0.6.1
Requires-Dist: platformdirs==4.10.0
Requires-Dist: pydantic==2.13.4
Requires-Dist: pydantic-settings==2.14.2
Requires-Dist: faster-whisper==1.2.1
Requires-Dist: numpy==2.2.6; python_version < '3.11'
Requires-Dist: numpy==2.4.6; python_version >= '3.11'
Requires-Dist: anthropic==0.116.0 ; extra == "all"
Requires-Dist: anthropic==0.116.0 ; extra == "anthropic"
Project-URL: Homepage, https://point53.ai/intercept.html
Provides-Extra: all
Provides-Extra: anthropic

# Point 53 Intercept™

**Your audio. Your transcript. Your machine.**

Point 53 Intercept records microphone and desktop audio simultaneously, transcribes it offline with faster-whisper (OpenAI Whisper weights, MIT), and runs the transcript through local LLMs via Ollama for summarization, analysis, and interactive chat. Ad-hoc evaluation of what was said and what it meant. Nothing leaves your machine unless you choose a cloud model. It is part of the Point 53 tools.

- **Local-first AI**: transcription via faster-whisper and summarization/analysis via [Ollama](https://ollama.com), all offline by default.
- **Mic + desktop capture**: records both sources mixed, or either one alone (`--mic-only` / `--desktop-only`).
- **Real-time query suggestions**: as you record, a fast local model proposes follow-up searches per chunk and prints them as clickable hyperlinks into your search engine (a local SearXNG instance by default).
- **Structured notes**: transcript (written to disk live, chunk by chunk), summary, and optional analysis as Markdown; no audio retained by default (opt-in `chunk_storage = "disk"` keeps a compact `audio.opus`).
- **Agent-ready**: ships an MCP server and a Skill, so agents can drive recording sessions through the same engine you do. See **Agent surface** below.

> **License:** [Mozilla Public License 2.0](https://mozilla.org/MPL/2.0/). See `LICENSE`, `NOTICE`, and `THIRD_PARTY_LICENSES.md`.

## Demo

> **WARNING:** The demo contains some instances of fast jump cuts and light changes. Do not watch if you have photosensitive epilepsy.
>
> **NOTE:** Ctrl + click (Cmd + click on macOS) to open the video in a new tab.

[![Point 53 Intercept™ Public Alpha Demo: Capture, Transcribe, Analyze](https://img.youtube.com/vi/XUXtzFht4kk/maxresdefault.jpg)](https://youtu.be/XUXtzFht4kk)

## Installation

Intercept installs as a standalone CLI tool with [uv](https://docs.astral.sh/uv/) from the Point 53 PEP 503 index. (PyPI hosts fail-loud stub packages only, to prevent name squatting; install from `dist.point53.ai`, not PyPI.)

**Configuration is the hard part for new users.** Read **[INSTALL.md](INSTALL.md)** for the full per-platform walkthrough (system dependencies, audio routing, device indices). Desktop-audio capture needs a one-time per-OS step (BlackHole on macOS, Stereo Mix or VB-Cable on Windows; Linux works out of the box via PulseAudio/PipeWire). Essentials are below.

### Linux / macOS

```bash
# Install uv if you don't have it
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install Point 53 Intercept
uv tool install p53-intercept --index https://dist.point53.ai/simple/
```

### Windows

```powershell
# Install uv if you don't have it
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

# Install Point 53 Intercept
uv tool install p53-intercept --index https://dist.point53.ai/simple/
```

To upgrade: `uv tool upgrade p53-intercept`. To uninstall: `uv tool uninstall p53-intercept`.

### Optional extras

Cloud (Anthropic) inference is an opt-in extra; the base install is local-only by design:

```bash
uv tool install 'p53-intercept[anthropic]' --index https://dist.point53.ai/simple/
```

## First Run

**Requirements**

- **Python 3.10+**: uv manages this for you.
- **ffmpeg**: records and mixes mic + system audio.
  - Linux: `sudo pacman -S ffmpeg libpulse` · `sudo apt install ffmpeg pulseaudio-utils` · `sudo dnf install ffmpeg pulseaudio-utils`
  - macOS: `brew install ffmpeg`
  - Windows: `winget install ffmpeg`
- **faster-whisper model**: no manual download; set `ai.whisper_model` in `capture.toml` (default `base`) and the int8 weights auto-download to `~/.local/share/point53/intercept/models/` on first run. Whisper weights are MIT-licensed at every size.
- **Ollama** running locally or on your network, with the default models pulled:
  ```bash
  ollama pull granite4:1b        # query-suggest (real-time, per-chunk)
  ollama pull gemma4:e4b         # summary + analysis
  ollama pull qwen3.5:9b         # chat
  ```

**Get going**

There is no `intercept config doctor` subcommand. The first command that loads config writes a fully-commented default `capture.toml` under `~/.config/point53/intercept/` and exits so you can edit before recording. After a `uv tool install`, trigger that once, edit the file, then check the install:

```bash
intercept record --name setup-check   # writes capture.toml and exits on first run
# edit ~/.config/point53/intercept/capture.toml
intercept doctor      # checks ffmpeg, audio backend (lists devices), whisper, config, LLM provider
```

Once `intercept doctor` is green, record for real (see Quick Start). On Linux, desktop audio works out of the box; on macOS and Windows, set up a loopback device first. See **[INSTALL.md](INSTALL.md)**.

A later suite host, **Point 53 Handler**, will orchestrate multi-tool workflows. Each tool stays independently installable.

## Quick Start

```bash
intercept record --name "meeting-2026-03-08" --title "Standup" --author "Me"
intercept record --name "lecture" --title "History 101" --author "Prof. Smith"
intercept record --name "call" --analysis-prompt "Extract all action items and decisions made"
intercept record --name "interview" --chat
intercept record --name "standup" --nofile          # print only, no session directory
intercept record --name "podcast" --desktop-only    # capture only system audio
intercept record --name "voicenote" --mic-only      # capture only the microphone
```

CLI `record` has no required duration: you stop interactively. MCP and Skill invocations use a timed `duration` argument instead (agents need a bounded session).

**During recording:** press `q` (or `Ctrl+C`, or `Ctrl+D`) to stop and proceed to AI post-processing, or `s` for an ad-hoc summary without stopping. A second `Ctrl+C` force-quits immediately. With `chunk_storage = "disk"`, audio left behind by a force-quit or crash is assembled automatically on the next run. With `--chat`, press `Ctrl+C` to exit the chat when finished.

### `record` options

| Option | Purpose |
| --- | --- |
| `--name` | Session label (becomes part of the session directory name) |
| `--title` | Title of recorded content (recorded in `manifest.json`) |
| `--author` | Creator / owner of the recorded audio |
| `--analysis-prompt` | Custom prompt for a secondary analysis pass over the transcript |
| `--chat` / `--chat-prompt` | Open an interactive chat after analysis (`--chat-prompt` seeds the first question) |
| `--nofile` | Skip text artifacts (transcript, notes, metadata); print only. Audio storage independently follows `chunk_storage` |
| `--desktop-only` / `--mic-only` | Record a single source (mutually exclusive; default records both mixed) |
| `--i-understand-the-risks` | Acknowledge a remote model will see this run's content (suppresses the cloud warning for one run) |

## Configuration

All runtime settings live in `~/.config/point53/intercept/capture.toml` (Pydantic-validated). A fully-commented default is written on first run. See **[INSTALL.md](INSTALL.md)** for platform-specific audio fields.

```toml
[audio]
chunk_duration = 70              # seconds per recording chunk
chunk_storage = "memory"         # memory | ephemeral | disk - see "Chunk storage"
mic_index = 0                    # macOS only (avfoundation index)
desk_index = -1                  # macOS only; -1 = mic-only

[ai]
provider = "ollama"              # default provider for all stages
ollama_host = "http://localhost:11434"
whisper_model = "base"           # tiny | base | small | medium | large-v3
query_url = "http://localhost:8888/search?q={query}"   # default: local SearXNG

[ai.query_suggest]               # real-time, per-chunk; small & fast
model = "granite4:1b"
ctx = 8192

[ai.summary]                     # post-recording summary
model = "gemma4:e4b"
ctx = 122880

[ai.analysis]                    # optional second pass (--analysis-prompt)
model = "gemma4:e4b"
ctx = 122880

[ai.chat]                        # optional interactive session (--chat)
model = "qwen3.5:9b"
ctx = 122880

[warnings]
cloud_model = "warn"             # "warn" or "silent"
chat_turn_limit = 10             # re-warn every N --chat turns (0 disables)
```

**Stages:** `query_suggest` (per-chunk real-time), `summary` (post-recording), `analysis` (optional second pass via `--analysis-prompt`), `chat` (optional interactive session). Each `[ai.<stage>]` can set its own `provider` to mix local and cloud (e.g. local `granite4:1b` for query-suggest, a cloud provider for analysis).

**API keys resolve env-first** (file-last): `P53_<PROVIDER>_API_KEY` > vendor env (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) > `capture.toml`. Routing any stage to a non-local `base_url` prints a cloud-usage warning; suppress with `--i-understand-the-risks` (one run) or `warnings.cloud_model = "silent"`.

### Chunk storage

`[audio].chunk_storage` chooses where in-flight audio chunks live while a session records: a confidentiality vs availability trade you set once:

| Mode | Where chunks live | Lifecycle |
| --- | --- | --- |
| `memory` (default) | Nowhere: ffmpeg streams raw PCM to the transcriber through a pipe | Audio never touches disk; a hard crash loses the session's audio |
| `ephemeral` | `~/.cache/point53/intercept/chunks/<session-id>/` | Each chunk is deleted right after transcription; leftovers from interrupted runs are swept, with a notice, at the next start |
| `disk` | `<session-dir>/chunks/`, assembled into `audio.opus` when recording stops | Audio is retained (~10 MB/hour opus vs ~115 MB/hour raw WAV); interrupted sessions are assembled on the next run |

`memory` is the most confidential: nothing to shred, snapshot, or sweep. `disk` is the most available: every exit (`q`, `Ctrl+C`, even a crash or power loss) converges on a single playable `audio.opus` (crash leftovers assemble at the next run; until then the raw chunks are ordinary WAVs, playable directly, e.g. `mpv <session-dir>/chunks/`). Retained audio lives in the session directory beside the manifest, so retention and legal-hold tooling governs it like any other artifact. `ephemeral` is the middle lane: useful for debugging capture problems, or as a fallback if the pipe path misbehaves on your platform; in practice it is barely more available than `memory`.

### Query suggestions

While you record, each transcribed chunk is sent to the fast `query_suggest` model, which proposes follow-up searches and prints them inline as clickable hyperlinks. Clicking one opens that query at whatever endpoint you point `[ai].query_url` at; `{query}` is the placeholder Intercept URL-encodes the suggested query into at run time.

```toml
# ~/.config/point53/intercept/capture.toml
[ai]
query_url = "http://localhost:8888/search?q={query}"   # default: local SearXNG
```

The default is a local [SearXNG](https://docs.searxng.org/) instance. Prefer SearXNG when you can: suggested queries stay on your network, you avoid commercial result pages that mix organic links with malvertising, and you are not depending on a commercial search engine's terms of service for how those links are opened. Nothing is fetched automatically: a suggestion only ever leaves your machine when you click it. Suggestions are for *your* browser; they are not a scrape path.

Platform-specific audio fields (`mic_index` / `desk_index` on macOS, `win_mic_device` / `win_desk_device` on Windows) are covered in **[INSTALL.md](INSTALL.md)**.

## Output

Each recording writes a self-describing session directory under `~/.local/share/point53/intercept/sessions/<session-id>/` (the id is `<timestamp>_<name>`):

- `transcript.md`: full transcript, written live, chunk by chunk, while recording, so a crash, kill, or orchestrator timeout never loses audio that was already transcribed
- `notes.md`: summary, plus analysis when `--analysis-prompt` is used (the primary artifact when present)
- `metadata.json`: run metadata (title, author, timings)
- `manifest.json`: typed manifest other tools read (`type: "intercept-session"`, schema v1.1, `ingestable_by: ["collector", "nightdesk"]`, plus the enterprise-coordination fields)
- `audio.opus`: the session's assembled audio (only with `chunk_storage = "disk"`)

No audio is retained by default: with `chunk_storage = "memory"` chunks stream from ffmpeg to the transcriber without ever touching disk, and `"ephemeral"` deletes its transient cache WAVs right after transcription. Only `"disk"` keeps audio, deliberately. There is no SQLite database; the Markdown files are the source of truth. Use `--nofile` to skip the text artifacts and print to the terminal only (audio storage still follows `chunk_storage`).

**Empty sessions are never dressed up.** A session that transcribed nothing is not summarized: no LLM stage runs against silence, and leaves no session directory behind (in `disk` mode, retained audio keeps a manifest). A capture that produces no audio at all (missing device, unreachable audio server) fails within seconds with an error naming the likely cause, instead of running out the clock.

## Agent surface - MCP + Skill

`intercept serve-mcp` starts a stdio MCP server over the same engine the CLI uses. Tool arguments are **flat, top-level schema properties** (the suite MCP convention); there is no nested params object:

| Tool | Purpose |
| --- | --- |
| `record` | Timed recording: capture for `duration` seconds, transcribe, summarize, optional `analysis_prompt` / `chat_question`. Args: `duration` (**required for MCP/Skill**, up to 86400 = 24 h), `name`, `title`, `author`, `mode` (`mic` \| `desk` \| `both`; default `both`) |
| `list_sessions` | Metadata for every recorded session |
| `get_session` | Metadata, manifest, and text artifacts for one session |
| `get_transcript` | Just the transcript: lighter than `get_session` |
| `run_doctor` | The `intercept doctor` checks, structured |
| `validate_config` | Validate `capture.toml`; resolved provider/model per stage |

`record` reports what actually happened: the result carries the `mode` recorded and `chunk_count`, a `warning` when a requested `both` fell back to mic-only because no desktop source was available, and an `error` (never a summary of silence) when nothing was transcribed. Timed sessions get the same live transcript writes and fail-fast capture as the CLI.

Orchestrators spawning `serve-mcp` must pass the user's session environment through to the subprocess (`XDG_RUNTIME_DIR` / `PULSE_*` on Linux): audio capture needs to reach the audio server. Point 53 Handler (planned) does this by default; if yours does not, the fail-fast error names what's missing.

For agent runtimes without MCP, a Skill (`SKILL.md`, shipped in-package) guides direct CLI use.

## File locations

Default paths by platform. On **macOS** and **Windows**, config and data resolve to the *same* directory.

**macOS with `XDG_*` set:** if `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, or `XDG_CACHE_HOME` is exported (common in dotfile-managed setups), Point 53 follows it instead of the `~/Library` default for that column, independently per variable. With the conventional values this collapses to the Linux/XDG layout, e.g. config at `~/.config/point53/intercept/`.

| Root | Linux | macOS | Windows |
| --- | --- | --- | --- |
| Config | `~/.config/point53/intercept/` | `~/Library/Application Support/point53/intercept/` | `%LOCALAPPDATA%\point53\intercept\` |
| Data | `~/.local/share/point53/intercept/` | `~/Library/Application Support/point53/intercept/` | `%LOCALAPPDATA%\point53\intercept\` |
| Cache | `~/.cache/point53/intercept/` | `~/Library/Caches/point53/intercept/` | `%LOCALAPPDATA%\point53\Cache\intercept\` |

Config holds `capture.toml`; data holds recorded `sessions/` and the downloaded faster-whisper `models/`; cache holds transient audio chunks (`ephemeral` chunk-storage mode only; discarded after transcription).

Operational logs are shared across the suite under the state root: `~/.local/state/point53/logs/` (`~/Library/Application Support/point53/logs/` on macOS, `%LOCALAPPDATA%\point53\logs\` on Windows).

Paths shown elsewhere in this README use the Linux/XDG form.

## TODO

**Planned**

- [ ] OS-keyring-backed API keys (beta-2): drop the plaintext key option from `capture.toml`
- [ ] Token / cost accounting for cloud model stages
- [ ] Pause control (beta-2): halt chunk iteration mid-`record` without ending the session, then resume
- [ ] Alternative transcription backends (beta-2): support non-Whisper speech-to-text models (NVIDIA and others) beyond faster-whisper
- [ ] Improved event emitting via MCP: the MCP path sets invocation provenance but skips the run-lifecycle events the CLI emits, so MCP-invoked sessions don't log the full process and interfere with Handler functionality

**Known limitations (alpha)**

- Desktop-audio capture needs a one-time loopback setup on macOS (BlackHole) and Windows (Stereo Mix / VB-Cable); Linux works out of the box. See [INSTALL.md](INSTALL.md).
- No speaker diarization: the transcript isn't split or labeled by speaker.

## License and Attribution

- **Point 53 Intercept source code:** Mozilla Public License, v. 2.0. See `LICENSE`.
- **Third-party runtime dependencies and external tools:** see `NOTICE` for a summary and `THIRD_PARTY_LICENSES.md` for full attribution.
- **Whisper and Ollama models:** not distributed with Intercept; each carries its own license from its respective publisher (Whisper weights are MIT; the default Ollama models are Apache-2.0).
- **Trademarks:** "Point 53" and "Point 53 Intercept" are trademarks of Point 53, LLC. MPL-2.0 section 2.3 excludes trademark rights from the copyright/patent grant; nothing in the license authorizes use of these marks. The CLI command `intercept` is a functional identifier, not a trademark claim on the generic English word.

