Metadata-Version: 2.4
Name: p53-collector
Version: 0.1.4
Summary: Structured retrieval, summarization, and classification of data through RSS feeds and web pages
Author-email: "Point 53, LLC" <dev@point53.ai>
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-Expression: MPL-2.0
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: End Users/Desktop
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Internet :: WWW/HTTP
Classifier: Topic :: Text Processing :: Markup :: Markdown
License-File: CONTRIBUTING.md
License-File: LICENSE
License-File: NOTICE
License-File: THIRD_PARTY_LICENSES.md
License-File: TRADEMARKS.md
Requires-Dist: click==8.4.2
Requires-Dist: dateparser==1.3.0
Requires-Dist: feedparser==6.0.12
Requires-Dist: httpx==0.28.1
Requires-Dist: mcp==1.28.1
Requires-Dist: ollama==0.6.1
Requires-Dist: platformdirs==4.10.0
Requires-Dist: protobuf==7.34.0
Requires-Dist: pydantic==2.13.4
Requires-Dist: selenium==4.41.0
Requires-Dist: tomli==2.4.1; python_version < '3.11'
Requires-Dist: webdriver-manager==4.0.2
Requires-Dist: anthropic==0.116.0 ; extra == "all"
Requires-Dist: pdfplumber==0.11.9 ; extra == "all"
Requires-Dist: anthropic==0.116.0 ; extra == "anthropic"
Requires-Dist: pdfplumber==0.11.9 ; extra == "pdf"
Project-URL: Homepage, https://point53.ai/collector.html
Provides-Extra: all
Provides-Extra: anthropic
Provides-Extra: pdf

# Point 53 Collector™

**Your sources. Your summaries. Your machine.**

Point 53 Collector pulls articles from RSS feeds and web pages, summarizes them with local AI, stores everything as Markdown (backed by a derived, rebuildable SQLite index), and prints curated digests on demand. No cloud accounts. No data leaves your network unless you choose a cloud model. It is part of the Point 53 tools.

- **Local-first AI**: summaries run on your hardware via [Ollama](https://ollama.com); your reading habits and the content itself never touch a third-party server by default.
- **Vendor flexibility**: swap Ollama models freely, or route a role to Anthropic or any OpenAI-compatible endpoint when it makes sense.
- **Structured, greppable output**: every article is categorized, timestamped, and queryable; filter by date, category, or keyword to Markdown or stdout.
- **Chat with your news**: after a digest, open an interactive chat about the articles using the same local models.
- **PDF-aware (opt-in)**: with the `[pdf]` extra and `pdf.enabled = true`, Collector downloads, scans, and summarizes linked PDFs from feeds like arXiv.

> **License:** [Mozilla Public License 2.0](https://mozilla.org/MPL/2.0/). See `LICENSE`, `NOTICE`, and `THIRD_PARTY_LICENSES.md`.

## Demo

> **WARNING:** The demo contains some instances of fast jump cuts and light changes. Do not watch if you have photosensitive epilepsy.
>
> **NOTE:** Ctrl + click (Cmd + click on macOS) to open the video in a new tab.

[![Point 53 Collector™ Public Alpha Demo: Aggregate, Summarize, Distill](https://img.youtube.com/vi/cfIJrPWGMIQ/maxresdefault.jpg)](https://youtu.be/cfIJrPWGMIQ)

## Installation

Collector installs as a standalone CLI tool with [uv](https://docs.astral.sh/uv/) from the Point 53 PEP 503 index. (PyPI hosts fail-loud stub packages only, to prevent name squatting; install from `dist.point53.ai`, not PyPI.)

### Linux / macOS

```bash
# Install uv if you don't have it
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install Point 53 Collector
uv tool install p53-collector --index https://dist.point53.ai/simple/
```

### Windows

```powershell
# Install uv if you don't have it
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

# Install Point 53 Collector
uv tool install p53-collector --index https://dist.point53.ai/simple/
```

### Optional extras

Cloud (Anthropic) inference and PDF feeds are opt-in extras; the base install is local-only by design:

```bash
uv tool install 'p53-collector[all]' --index https://dist.point53.ai/simple/
# or pick one: [anthropic] (cloud SDK) · [pdf] (linked-PDF summarization)
```

## First Run

**Requirements**

- **Python 3.10+**: uv manages this for you.
- **Ollama** running locally or on your network, with the default models pulled:
  ```bash
  ollama pull gemma4:e4b         # article summarization
  ollama pull qwen3.5:9b         # distill --chat
  ollama pull granite4.1:8b      # SEARCH-page links → RSS
  ```
- **Firefox**: the default scraping browser. Selenium drives a separate instance to keep scraping out of your everyday browser's profile, and Collector assumes that everyday browser is Chrome. Adjust the [configuration](#configuration) if Firefox is your default, so Chrome (or `undetected-chrome`) runs for Collector instead.

**Get going**

`uv tool install` does not create config files. After install, create defaults and check the install:

```bash
collector config doctor
collector doctor

# Verify the CLI is on your PATH
collector --help
```

`collector config doctor` writes two TOML files under `~/.config/point53/collector/` (`feeds.toml`, `briefings.toml`) with sensible defaults. Edit them to add your sources and point at your Ollama endpoint, then:

```bash
collector update      # fetch + summarize new articles
collector distill     # write a Markdown digest of unread articles
```

That's the whole base setup. Everything below gets extensible.

A later suite host, **Point 53 Handler**, will orchestrate multi-tool workflows. Each tool stays independently installable.

## Configuration

Two TOML files under `~/.config/point53/collector/`. Run `collector config doctor` to create them and `collector config validate` to check them against the schema.

### `feeds.toml`: your sources

```toml
blocklist = []                    # URLs to skip during updates

[[feeds]]
name     = "Example Feed"
category = "Technology"           # free-form label; filter on it with `distill --category`
type     = "RSS"                  # "RSS" (parsed directly) or "SEARCH" (scraped + LLM-curated)
link     = "https://example.com/rss"
note     = ""                     # optional scraping hint
scroll   = 0                      # SEARCH pages: extra scrolls before harvesting links
```

### `briefings.toml`: models & operational settings

| Section | Purpose |
| --- | --- |
| `[webdriver]` | Browser (default `firefox`; also `chrome`, `undetected-chrome`), `page_load_timeout`, `ignore_certificate_errors`, optional `profile_dir` |
| `[pdf]` | `enabled` + `max_size_mb` for downloading and summarizing linked PDFs |
| `[models.article_summary]` | LLM role that summarizes each scraped page |
| `[models.search_to_rss]` | LLM role that turns SEARCH-page links into structured RSS |
| `[models.article_chat]` | LLM role backing `distill --chat` |
| `[rate_limits]` | `summary_sleep` / `search_sleep` throttles between requests |
| `[timeouts]` | `llm` request timeout (seconds) |
| `[warnings]` | `cloud_model` (`"warn"`/`"silent"`) and `chat_turn_limit` (re-warn every N turns of `distill --chat`; `0` disables) |
| `[profile]` | browser session `default`: `"sandboxed"` (ephemeral) or `"credentialed"` (persistent) |

Each `[models.<role>]` takes a `provider` (`ollama` / `anthropic` / `openai-compatible`), `model`, `base_url`, `context_window`, and an optional `api_key`; mix providers per role freely:

```toml
[models.article_summary]
provider       = "ollama"
model          = "gemma4:e4b"
base_url       = "http://localhost:11434"
context_window = 122880
```

**API keys resolve env-first** (file-last): `P53_<PROVIDER>_API_KEY` > vendor env (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) > the literal `api_key` in `briefings.toml`. Local Ollama / LM Studio need no key. Routing any role to a non-local `base_url` prints a one-line cloud-usage warning on stderr; suppress with `warnings.cloud_model = "silent"` or `--i-understand-the-risks`.

## How it works (risk reduction)

```
feeds.toml ─▶ RSS parse / SEARCH scrape ─▶ LLM curate (SEARCH only)
                      │
                      ▼
        check_fetch_url (http/https only; private hosts allowed with notice)
                      │
                      ▼
              scrape page / PDF ─▶ LLM summarize ─▶ briefing .md + index
                      │
                      ▼
                 distill ─▶ Markdown digest (+ optional chat)
```

Fetch targets are vetted before navigation:

- **Schemes:** only `http` and `https` are fetched; other schemes are refused.
- **Private / loopback / `.local`:** allowed for intranet research; Collector prints a transparency notice so you know a private host is in play.
- **Terminal output:** control characters are stripped via `_strip_ctl` before untrusted titles, links, or model text are printed, which blocks OSC/ANSI injection from scraped or model content.

## File locations

Default paths by platform. On **macOS** and **Windows**, config and data resolve to the *same* directory.

**macOS with `XDG_*` set:** if `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, or `XDG_CACHE_HOME` is exported (common in dotfile-managed setups), Point 53 follows it instead of the `~/Library` default for that column, independently per variable. With the conventional values this collapses to the Linux/XDG layout, e.g. config at `~/.config/point53/collector/`.

| Root | Linux | macOS | Windows |
| --- | --- | --- | --- |
| Config | `~/.config/point53/collector/` | `~/Library/Application Support/point53/collector/` | `%LOCALAPPDATA%\point53\collector\` |
| Data | `~/.local/share/point53/collector/` | `~/Library/Application Support/point53/collector/` | `%LOCALAPPDATA%\point53\collector\` |
| Cache | `~/.cache/point53/collector/` | `~/Library/Caches/point53/collector/` | `%LOCALAPPDATA%\point53\Cache\collector\` |

Config holds `feeds.toml` and `briefings.toml`; data holds the per-run briefings, the derived `index.db`, distillates, and browser profiles. The Markdown briefings are the source of truth; `index.db` is rebuildable with `collector reindex`.

Operational logs are shared across the suite under the state root: `~/.local/state/point53/logs/` (`~/Library/Application Support/point53/logs/` on macOS, `%LOCALAPPDATA%\point53\logs\` on Windows).

Paths shown elsewhere in this README use the Linux/XDG form.

## Command Reference

```bash
# Fetch and summarize new articles (headless + verbose if you like)
collector update --headless --verbose

# Generate a Markdown digest of unread articles
collector distill

# Filter by category and full-text search (across read + unread)
collector distill --category "AI" --string-search "open weights" --include-seen

# Restrict to a date range (revisit a past window)
collector distill --start 01-01-2026 --end 02-01-2026 --include-seen

# Chat about what you've read
collector distill --chat "What were the most significant open-model releases this week?" --include-seen

# Rebuild the derived index from the Markdown source-of-truth
collector reindex

# Health checks and config management
collector doctor
collector config validate
collector config doctor

# Manage authenticated browser sessions (see Authenticated sessions)
collector profile status

# Serve Collector's tools to agents over MCP (stdio)
collector serve-mcp
```

> **`distill` acts on new (un-distilled) summaries by default.** A plain `collector distill` consumes the unread queue and marks those runs seen; filtered queries (`--category`, `--string-search`, `--start`/`--end`) are read-only and never consume. Add `--include-seen` to search or revisit articles you've already distilled.

Markdown briefings under `~/.local/share/point53/collector/briefings/<run>/` are the source of truth; `index.db` is a derived FTS index, rebuildable any time with `collector reindex`. Each run directory includes a `manifest.json` (schema v1.1) for cross-tool coordination.

## Authenticated sessions

Some sources sit behind a login. `collector profile` manages a persistent **credentialed** browser profile so Collector can scrape pages you're signed into, kept separate from the default ephemeral **sandboxed** session.

```bash
# Open a visible browser on the credentialed profile; log in, then press Enter to save
collector profile init                 # optional: --browser chrome | firefox

# Show the credentialed profile's path, size, and last-modified time
collector profile status

# Delete the credentialed profile and every saved session (prompts to confirm)
collector profile reset
```

Log into your sites with **"Remember me"** enabled during `profile init`, then point updates at that profile:

```bash
collector update --profile credentialed
```

…or make it the default in `briefings.toml`:

```toml
[profile]
default = "credentialed"               # default is "sandboxed" (ephemeral)
```

> **Agents never get your logins.** Credentialed profiles work only through the `collector` CLI; the MCP server is sandboxed-only, so an agent can't drive your authenticated sessions.

## Scheduling

Collector is a stateless CLI. No daemon. Schedule unattended updates with your OS scheduler. The repo ships `cron.sh` as a starting point:

```bash
#!/bin/bash
export PATH="$HOME/.local/bin:$PATH"
cd /path/to/p53_collector
uv run collector update --headless
```

Point a cron entry at it (adapt the path and cadence to your setup):

```cron
# Fetch + summarize every weekday at 6am
0 6 * * 1-5  /path/to/p53_collector/cron.sh >> "$HOME/collector-cron.log" 2>&1
```

If you installed Collector with `uv tool install`, `collector` is already on your `PATH`; a cron line can call `collector update --headless` directly, no wrapper needed. Read the queue whenever with `collector distill`.

## Agents (MCP & Skills)

Collector exposes its capabilities to agents two ways.

**MCP**: `collector serve-mcp` runs a stdio MCP server so agents (or the planned Point 53 Handler) can drive Collector programmatically. Five tools:

| Tool | Purpose |
| --- | --- |
| `update` | Fetch, scrape, and summarize new articles (sandboxed sessions only) |
| `distill` | Query stored articles and return Markdown; filtered queries are read-only peeks |
| `reindex` | Rebuild the derived index from the Markdown source-of-truth |
| `list_runs` | List update runs with their manifest metadata |
| `get_article` | Read a specific article's frontmatter and body |

**Skills**: for runtimes without MCP, the bundled `skills/collector/SKILL.md` documents the same surface for direct CLI invocation.

> The MCP server is **sandboxed-only**; credentialed browser profiles are never exposed to agents (see [Authenticated sessions](#authenticated-sessions)).

## TODO

**Planned**

- [ ] OS-keyring-backed API keys (beta-2): drop the plaintext key option from `briefings.toml`
- [ ] Token / cost accounting for cloud model roles
- [ ] Fix `collector doctor` not detecting installed web browsers on macOS and Windows
- [ ] Fix mismatched index counts after `collector reindex`
- [ ] Document the engine API and wiring instructions for embedding Collector as a library

**Known limitations (alpha)**

- `SEARCH`-type feeds rely on a live browser (Selenium + Firefox/Chrome) plus LLM curation, so they're slower and more fragile than plain RSS.
- PDF summarization is opt-in (`[pdf]` extra); large or scanned PDFs may be skipped per `pdf.max_size_mb`.

## License and Attribution

Point 53 Collector is licensed under the [Mozilla Public License 2.0](https://mozilla.org/MPL/2.0/).

- `LICENSE`: full MPL-2.0 text
- `NOTICE`: copyright and trademark notice, third-party summary
- `THIRD_PARTY_LICENSES.md`: per-dependency attribution and obligations

Ollama, Anthropic, and other provider models are not distributed with Collector; each carries its own license from its respective publisher.

"Point 53" and "Point 53 Collector" are trademarks of Point 53, LLC. MPL-2.0 section 2.3 excludes trademark rights from the copyright/patent grant; nothing in the license authorizes use of these marks. The bare `collector` CLI name is a functional identifier, not a trademark claim on the generic English word.

