Metadata-Version: 2.4
Name: p53-nightdesk
Version: 0.1.4
Summary: Point 53 Nightdesk — iterative metric optimization with AI-assisted hypothesis generation.
Keywords: optimization,autoresearch,ollama,anthropic,local-llm,point53
Author-email: "Point 53, LLC" <dev@point53.ai>
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-Expression: MPL-2.0
Classifier: Development Status :: 3 - Alpha
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: POSIX :: Linux
Classifier: Operating System :: MacOS
Classifier: Operating System :: Microsoft :: Windows
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
License-File: CONTRIBUTING.md
License-File: LICENSE
License-File: NOTICE
License-File: THIRD_PARTY_LICENSES.md
License-File: TRADEMARKS.md
Requires-Dist: rich==15.0.0
Requires-Dist: httpx==0.28.1
Requires-Dist: mcp==1.28.1
Requires-Dist: platformdirs==4.10.0
Requires-Dist: pydantic==2.13.4
Requires-Dist: tomli==2.4.1; python_version < '3.11'
Requires-Dist: anthropic==0.116.0 ; extra == "all"
Requires-Dist: anthropic==0.116.0 ; extra == "anthropic"
Project-URL: Homepage, https://point53.ai/nightdesk.html
Provides-Extra: all
Provides-Extra: anthropic

# Point 53 Nightdesk™

**An autonomous tuner for anything with a number to move.**

Point 53 Nightdesk is a flag-based CLI for AI-driven iterative metric optimization. Describe what you want to optimize; Nightdesk scaffolds the project, runs experiments, tracks scores, keeps improvements, reverts regressions, and accumulates knowledge, across as many projects as you want. If there's a number you want to maximize or minimize and a way to measure it from a shell command, Nightdesk can loop on it: ML training, trading strategies, energy efficiency, prompt engineering, game balance, compiler flags, API latency, infrastructure cost, scientific parameter tuning, and more. It is part of the Point 53 tools.

> **Warnings (read before you loop)**
>
> 1. **`shell=True`.** Loop and scoring processes run shell commands with `shell=True`. Only use AI providers you trust, and review generated scripts and output before running baselines or loops.
> 2. **Courier during loops.** Because of (1), enabling Courier inside the loop is a feature lift at a security cost. Prefer Courier on `new` and `edit` only. Keep `[courier].loop = "off"` unless you deliberately accept the risk.
> 3. **Experimental frontier.** Nightdesk is the most frontier-oriented and most experimental tool in the Point 53 suite. Use at your own risk for applications and environments that matter.

- **Seven commands, zero memorization**: `new` / `list` / `edit` / `test` / `loop` / `split` / `help`; every command except `new` takes an explicit project name. No active-project state. Scriptable, agent-friendly, overnight-safe.
- **Local-first AI**: three independently configurable LLM roles; the defaults run fully local on [Ollama](https://ollama.com).
- **Keeps what works**: auto-commits improvements, reverts regressions, and enforces guardrails against degenerate optimization.
- **Accumulates knowledge**: dead ends and insights persist per project, so the AI never re-explores a dead end.
- **Optional web research**: when Point 53 Courier is installed, stages you enable can gather live web context on demand.

```
nightdesk new                    Create a project from a description (AI-scaffolded)
nightdesk list                   Show all projects
nightdesk edit <name>            Chat with AI, configure, or manage lifecycle
nightdesk test <name>            Verify the project's scoring command works
nightdesk loop <name>            Run the autonomous optimization loop
nightdesk split <src> <new>      Fork a project as a new variant
nightdesk help                   Show the menu
```

> **License:** [Mozilla Public License 2.0](https://mozilla.org/MPL/2.0/). See `LICENSE`, `NOTICE`, and `THIRD_PARTY_LICENSES.md`.

## Demo

> **WARNING:** The demo contains some instances of fast jump cuts and light changes. Do not watch if you have photosensitive epilepsy.
>
> **NOTE:** Ctrl + click (Cmd + click on macOS) to open the video in a new tab.

[![Point 53 Nightdesk™ Public Alpha Demo: Hypothesize, Iterate, Synthesize](https://img.youtube.com/vi/dQV_Xqdvi7Y/maxresdefault.jpg)](https://youtu.be/dQV_Xqdvi7Y)

## Installation

Nightdesk installs as a standalone CLI tool with [uv](https://docs.astral.sh/uv/) from the Point 53 PEP 503 index. (PyPI hosts fail-loud stub packages only, to prevent name squatting; install from `dist.point53.ai`, not PyPI.)

### Linux / macOS

```bash
# Install uv if you don't have it
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install Point 53 Nightdesk
uv tool install p53-nightdesk --index https://dist.point53.ai/simple/
```

### Windows

```powershell
# Install uv if you don't have it
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

# Install Point 53 Nightdesk
uv tool install p53-nightdesk --index https://dist.point53.ai/simple/
```

### Optional extras

Cloud (Anthropic) inference is an opt-in extra; the base install is local-only by design:

```bash
uv tool install 'p53-nightdesk[all]' --index https://dist.point53.ai/simple/
# from a local clone (development): uv tool install -e ".[all]"
```

## First Run

**Requirements**

- **Python 3.10+**, **uv**, and **git**. Git is load-bearing: Nightdesk auto-commits each kept experiment and reverts regressions.
- **Ollama** running locally or on your network, with the default models pulled:
  ```bash
  ollama pull qwen3.5:9b        # new + edit roles
  ollama pull gemma4:e4b        # loop role
  ```

Nightdesk drives three independently configurable LLM roles: **`new`** (scaffolding), **`edit`** (chat), and **`loop`** (autonomous iteration). The defaults are Ollama for all three, so a fresh install runs fully local. A typical upgrade pairs a stronger cloud model for `new`/`edit` with Ollama for `loop` to keep cost down across long loops. More capable models often improve scores faster, at greater expense. Roles live in `~/.config/point53/nightdesk/research.toml`.

**Get going**

`uv tool install` does not always leave a config in place the way the install script does. After install, seed defaults and check the environment:

```bash
nightdesk config doctor   # writes research.toml + palette defaults if missing
nightdesk doctor

# 1. Create a project. Interactive: gather details, AI generates files, you iterate
nightdesk new "optimize Python web API response latency with locust"

# 2. Verify the scoring command actually runs
nightdesk test my-api-project

# 3. Let the AI drive the loop
nightdesk loop my-api-project --max-iterations 20
```

You'll see each iteration stream by: the AI's hypothesis, the file edit, the experiment run, the score, keep or revert.

A later suite host, **Point 53 Handler**, will orchestrate multi-tool workflows. Each tool stays independently installable.

## Quick Start

```bash
# Graceful stop of a running loop (preferred)
nightdesk loop my-api-project --stop        # halt after the current experiment

# Detach for overnight runs
nightdesk loop my-api-project --detach --outfile ~/api.log
```

Wake up to a log of experiments and a better score. Loops are stateful while running: another terminal can still talk to the same project (`edit`, `list --name`, `loop --stop`). Prefer `loop --stop` for a clean halt. **Ctrl+C is not a graceful exit**; it can interrupt mid-commit or mid-experiment and leave the project in a partial state.

### Managing projects

```bash
nightdesk list                              # active projects, stacked format
nightdesk list --name my-api-project        # detailed view: experiments, knowledge, loop history
nightdesk list --all                        # include hidden/archived projects

# Lifecycle (all via edit)
nightdesk edit my-project --hide            # soft-delete (keep files)
nightdesk edit my-project --show            # restore from hidden
nightdesk edit my-project --delete          # hard-delete (wipe files)

# Fork to try a different approach without losing progress
nightdesk split trading-v1 trading-v2-aggressive
```

### Reconfiguring the LLM (takes effect immediately)

`--role` selects which role the next provider/model/url flags apply to; it defaults to `loop`:

```bash
nightdesk edit my-project --model qwen3.5:9b                                   # loop role (default)
nightdesk edit my-project --role edit --provider anthropic --model claude-opus-4-7
nightdesk edit my-project --role loop --url http://10.0.0.3:11434             # different Ollama server
# To change the `new` role, edit ~/.config/point53/nightdesk/research.toml directly:
# no project exists at scaffold time, so it's global-only.
```

### Chatting with the AI about a project

```bash
nightdesk edit my-project
```

Drops you into a chat session with the project's configured AI, which has full context: `program.md`, tunable files, results history, knowledge base. Ask about results, request file edits (the AI proposes, you approve), record insights or dead-ends. The AI uses structured markers: `EDIT: <filename>` + code block (wrapper asks to apply), `INSIGHT: <text>` and `DEAD-END: <text>` (auto-recorded in `knowledge.json`). Type `done`, `quit`, or press **Ctrl+D** to exit.

## Configuration

Global config lives in `~/.config/point53/nightdesk/research.toml` (Pydantic-validated), created on first use or via `nightdesk config doctor`. Per-project config is snapshotted into `<project>/.nightdesk/config.json` at `nightdesk new` time, so editing `research.toml` later affects only future scaffolds; existing projects stay hermetic.

### LLM roles

Three roles, configured under `[llm.new]`, `[llm.edit]`, `[llm.loop]`:

- **`new`**: `nightdesk new` scaffolding (file generation + optional Courier web research). Capability helps; runs once per project. Default `qwen3.5:9b`.
- **`edit`**: `nightdesk edit` chat (assess issues, propose file edits). Capability matters here too. Default `qwen3.5:9b`.
- **`loop`**: `nightdesk loop` autonomous iteration. Runs many times; cheap/local is the right call to keep cost down across long loops. More capable models often improve scores faster, at greater expense. Default `gemma4:e4b`.

The `loop` and `edit` roles can be changed per-project via `nightdesk edit <name> --role {edit|loop} --provider X --model Y`. The `new` role is global-only: edit `[llm.new]` in `research.toml` directly.

**API keys resolve env-first** (file-last): `P53_<PROVIDER>_API_KEY` > vendor env (`ANTHROPIC_API_KEY`, `OPENAI_API_KEY`) > a literal `api_key` on `[llm.<role>]` in `research.toml`. The file key is read live and is never snapshotted into a project's `.nightdesk/config.json`. Routing a role to a non-local `base_url` prints a cloud-usage warning.

```bash
# Cloud provider example: env is the recommended path
export ANTHROPIC_API_KEY=sk-ant-...
export P53_ANTHROPIC_API_KEY=sk-ant-...      # or, the cross-tool P53_ key (takes precedence)
```

### Web research (Courier)

If `p53-courier` is installed and `[courier].enabled = true` in `research.toml`, the LLM in any stage set to `on_demand` can emit `GATHER: <query>` lines mid-response to fetch live web context. The system pauses, runs Courier, and resumes with the results appended, capped at `[courier].max_rounds` per turn.

```toml
[courier]
enabled = true                # master kill switch: overrides everything below
new  = "on_demand"            # GATHER: during scaffolding (recommended when Courier is on)
edit = "on_demand"            # GATHER: in nightdesk edit chat
loop = "on_demand"            # current default; strongly prefer "off" (see Warnings)
# Set any stage to "off" to disable Courier there without affecting other stages.
```

Current code defaults are on-demand for all three stages. Because loop experiments execute shell commands, set `loop = "off"` unless you intentionally want live web data inside the autonomous loop and accept the security cost.

### Package palette

A global allowlist/blocklist constrains which Python packages the AI may build projects with. One package per line, `#` for comments, created by `nightdesk config doctor`:

```
~/.config/point53/nightdesk/palette/
  allowlist.txt   # preferred packages (defaults: httpx, rich, pytest, numpy, pandas, scikit-learn)
  blocklist.txt   # forbidden packages (default: empty — see TODO)
```

Enforcement is layered. The scaffold LLM's system prompt marks the palette STRICT: prefer allowlisted packages, never emit blocklisted ones. The declared dependencies in the generated `requirements.txt` / `pyproject.toml` are then hard-checked twice — before `nightdesk new` writes anything to disk, and again before `uv sync` installs (which also covers projects reworked through `nightdesk edit`). A blocklist hit aborts with a palette violation; a package outside a non-empty allowlist is warned as off-palette but allowed.

The hard checks read **declared third-party dependencies**. Standard-library modules (`os`, `subprocess`, …) are never declared as dependencies, so blocklisting them today steers the scaffold prompt but is not mechanically enforced — see TODO.

## Use Cases

**Machine Learning**. Tune hyperparameters, architecture choices, training schedules. Score: validation loss, accuracy, F1.

**Trading Strategies**. Optimize entry/exit logic, position sizing, risk parameters. Score: PnL, Sharpe ratio, win rate.

**Energy Systems**. Tune HVAC schedules, battery dispatch, and solar storage strategies. Score: kWh cost, round-trip efficiency, peak-demand reduction.

**Prompt Engineering**. Iterate on system prompts, few-shot examples, output formats. Score: task accuracy, response quality metric.

**Game Balance**. Tune difficulty curves, economy parameters, spawn rates. Score: simulated completion rate, player satisfaction proxy.

**Infrastructure**. Optimize cloud configs, caching strategies, database tuning. Score: latency p99, cost per request, throughput.

**Scientific Research**. Optimize reaction parameters, simulation configs, experimental conditions. Score: yield, efficiency, convergence rate.

**Compiler/Build Optimization**. Tune compiler flags, build configurations, link-time optimizations. Score: binary size, execution time, compile time.

## Command Reference

### `nightdesk new`
Interactive: gather details (scoring method, constraints, acceptance criteria), AI generates all files (`program.md`, scoring scripts, configs, `requirements.txt`), you iterate until satisfied. Auto-installs Python deps into a project-local `.venv`.

### `nightdesk list [flags]`
| Flag | Effect |
|---|---|
| *(none)* | Active projects, stacked format |
| `--all` | Include hidden/archived (adds Status row) |
| `--name <name>` | Single project, detailed view (experiments, knowledge, loop runs) |

### `nightdesk edit <name> [flags]`
| Flag | Effect |
|---|---|
| *(none)* | Enter AI chat session, uses the project's `edit` role (Ctrl+D to exit) |
| `--hide` | Soft-delete: status → archived, files kept |
| `--show` | Un-hide a hidden project |
| `--delete` | Hard-delete: unregister + wipe files |
| `--role <edit\|loop>` | Which LLM role the next provider/model/url flags apply to (default: `loop`) |
| `--provider <ollama\|anthropic\|openai-compatible>` | Change LLM provider for the chosen role |
| `--model <name>` | Change LLM model for the chosen role |
| `--url <url>` | Change provider endpoint URL for the chosen role |

The `new` role (used by `nightdesk new`) is global; edit `[llm.new]` in `~/.config/point53/nightdesk/research.toml` directly.

### `nightdesk test <name>`
Runs the project's scoring command, verifies the score parses correctly, shows diagnostic suggestions on failure. Does **not** record an experiment: this is a sanity check, not a commit.

### `nightdesk loop <name> [flags]`
| Flag | Effect |
|---|---|
| *(none)* | Foreground streaming until done or stopped |
| `--detach` | Run in background, log to `.nightdesk/loop.log` |
| `--outfile <path>` | Override log destination |
| `--stop` | Graceful halt (sentinel file); preferred over Ctrl+C |
| `--max-iterations N` | Default 50 |
| `--stop-plateau N` | Default 5: stop if last N experiments within 1% of best |

Auto-baselines on first run if no experiments exist yet.

### `nightdesk split <source> <new-name>`
Fork a project: copies files (skipping `.git`, `.venv`, `results.tsv`), creates a fresh git repo, registers as a new project. Use for parallel experiment tracks without losing the original.

## Concepts

**The optimization loop:** Hypothesis → edit → commit → run → measure → keep or revert → record → repeat. One variable at a time. No grid searches.

**Scaffolding:** AI generates everything needed for a loop from a plain English description. Iterative: review, give feedback, regenerate until right. Context is saved for later use.

**Registry:** SQLite database at `~/.local/share/point53/nightdesk/registry.db` tracks all projects and loop run history. It's an index; project directories are always the source of truth.

**Knowledge:** Dead ends and insights accumulate in each project's `knowledge.json`. The AI reads them before suggesting or generating hypotheses. Never re-explores a dead end.

**Guardrails:** Bounds on metrics that prevent degenerate optimization. Score-explosion detection (>100x previous best) plus configurable min/max metric limits. The autonomous loop auto-reverts on guardrail violations.

**Two-pane workflow:** Control plane (an agent runtime or human, thinking side) + execution plane (Nightdesk, bookkeeping side). For autonomous loops, both merge into one process.

## Project Layout

```
your-project/
├── .nightdesk/
│   ├── config.json           # Run command, score field, LLM config
│   ├── knowledge.json        # Dead ends + insights
│   ├── scaffold_context.txt  # Saved scaffold context
│   ├── loop.log              # Background loop output (--detach only)
│   └── loop.stop             # Sentinel file (created by loop --stop)
├── .venv/                    # Auto-created for Python projects
├── manifest.json             # Schema v1.1 project manifest
├── program.md                # Research directions (AI-generated)
├── results.tsv               # Experiment log
├── requirements.txt          # Python deps (auto-installed)
├── <tunable files>           # The knobs, edited by loop or chat
└── .gitignore
```

## File locations

Default paths by platform. On **macOS** and **Windows**, config and data resolve to the *same* directory.

**macOS with `XDG_*` set:** if `XDG_CONFIG_HOME`, `XDG_DATA_HOME`, or `XDG_CACHE_HOME` is exported (common in dotfile-managed setups), Point 53 follows it instead of the `~/Library` default for that column, independently per variable. With the conventional values this collapses to the Linux/XDG layout, e.g. config at `~/.config/point53/nightdesk/`.

| Root | Linux | macOS | Windows |
| --- | --- | --- | --- |
| Config | `~/.config/point53/nightdesk/` | `~/Library/Application Support/point53/nightdesk/` | `%LOCALAPPDATA%\point53\nightdesk\` |
| Data | `~/.local/share/point53/nightdesk/` | `~/Library/Application Support/point53/nightdesk/` | `%LOCALAPPDATA%\point53\nightdesk\` |

Config holds `research.toml` and `palette/`; data holds the SQLite `registry.db` and the default `projects/` directory. Per-project files (`program.md`, `results.tsv`, `knowledge.json`, tunables) live inside each project directory. See [Project Layout](#project-layout).

Operational logs are shared across the suite under the state root: `~/.local/state/point53/logs/` (`~/Library/Application Support/point53/logs/` on macOS, `%LOCALAPPDATA%\point53\logs\` on Windows).

Paths shown elsewhere in this README use the Linux/XDG form.

## Using With Agent Runtimes

**Let Nightdesk run the loop.** `nightdesk loop` is purpose-built for autonomous iteration and gives you guarantees a raw agent session won't: guaranteed forward progression until a defined stop point (max iterations, plateau, sentinel, or guardrail violation), scope guardrails that auto-revert regressions, and score-explosion detection. For the optimization run itself, prefer the loop.

Where an agent runtime fits naturally is the thinking and orchestration around it: reading project state, reasoning about results, and driving Nightdesk's commands from a split terminal:

```bash
nightdesk list
nightdesk list --name my-project
nightdesk test my-project
# (reason about results, propose changes)
nightdesk loop my-project --detach      # hand off to autonomous mode
```

### Agent chat in place of `nightdesk edit` (optional, less supported)

`nightdesk edit` drops you into a chat session with the project's configured AI and full project context. If you'd rather do that refinement work in a raw agent session, that's fine. Just understand it's an optional, less-supported path. It relies on the agent to understand the project layout and edit the right files; the project's on-disk context (`program.md`, the tunable files, `results.tsv`, `knowledge.json`) is usually rich enough for it to do so well, but the responsibility shifts to you and the tool. To work this way:

1. **Find the project directory**: `nightdesk list --name my-project` prints its `Path:`.
2. **Run the agent from that directory** so the whole project is in scope.
3. **Treat the session like `nightdesk edit`**: discuss results, edit tunable files, record insights and dead-ends, but extensible through your own workflows.

See `AGENTS.md` for architecture details and integration notes.

## TODO

**Planned**

- [ ] Pluggable research backend (`p53.nightdesk.backend`) (beta-2): a vertically integrated stack with Courier is where we start; the pluggable interface lands later
- [ ] OS-keyring-backed API keys (beta-2): drop the plaintext key option from `research.toml`
- [ ] Unify commandline structure and first-run naming conventions (beta-2)
- [ ] Improve `new` mechanics for TUI and MCP (beta-2)
- [ ] Token / cost accounting for cloud model roles
- [ ] Scoped Courier criteria for loop (beta-2): live-data benefit with narrow enough bounds to restrict prompt injection (see suite report)
- [ ] Sane defaults for the package-palette blocklist: ship a blocked-by-default set in `blocklist.txt` instead of the current empty file
- [ ] Import-level palette enforcement: extend the hard checks beyond declared dependencies to imports in generated and loop-edited code, so stdlib entries like `os` bind mechanically

**Known limitations (alpha)**

- Web research requires Point 53 Courier installed and `[courier].enabled = true`.
- Interactive (human) scoring can't run under `loop --detach`, and scoring scripts must be plain text + `input()` (no TUI / curses / cursor positioning).
- The registry is a SQLite index; project directories remain the source of truth.

## License and Attribution

- **Point 53 Nightdesk source code:** Mozilla Public License, v. 2.0. See `LICENSE`.
- **Third-party runtime dependencies and external tools:** see `NOTICE` for a summary and `THIRD_PARTY_LICENSES.md` for full attribution.
- **Ollama, Anthropic, and other provider models:** not distributed with Nightdesk; each carries its own license from its respective publisher.
- **Trademarks:** "Point 53" and "Point 53 Nightdesk" are trademarks of Point 53, LLC. MPL-2.0 section 2.3 excludes trademark rights from the copyright/patent grant; nothing in the license authorizes use of these marks. The CLI command `nightdesk` is a functional identifier, not a trademark claim on the generic English word.

