Scaffold monorepo structure for all pipeline components
Add top-level README, .gitignore, and per-component directories (recording, labeling backend/frontend, training, inference, control) with READMEs and pyproject.toml/package skeletons per CLAUDE.md §2-3. No implementation yet, just structure to build against. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
parent
e0996275ff
commit
65155ce22e
32
.gitignore
vendored
Normal file
32
.gitignore
vendored
Normal file
@ -0,0 +1,32 @@
|
||||
# Python
|
||||
.venv/
|
||||
venv/
|
||||
__pycache__/
|
||||
*.pyc
|
||||
*.pyo
|
||||
*.egg-info/
|
||||
build/
|
||||
dist/
|
||||
.pytest_cache/
|
||||
.mypy_cache/
|
||||
.ruff_cache/
|
||||
|
||||
# Data / models / recordings (large, not source)
|
||||
*.mp4
|
||||
*.mkv
|
||||
*.avi
|
||||
data/
|
||||
datasets/
|
||||
checkpoints/
|
||||
*.pt
|
||||
*.pth
|
||||
*.onnx
|
||||
|
||||
# OS
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
|
||||
# Editors
|
||||
.vscode/
|
||||
.idea/
|
||||
*.swp
|
||||
162
CLAUDE.md
Normal file
162
CLAUDE.md
Normal file
@ -0,0 +1,162 @@
|
||||
# SpelunkAI — Project Definition
|
||||
|
||||
- **Project name:** SpelunkAI
|
||||
- **Git repository:** https://git.raggonium.net/Jonas/spelunkai
|
||||
- **Commit convention:** Every commit message must include the name of the Claude model that made it (i.e. the current model name in use at the time of the commit).
|
||||
|
||||
## 1. Project Goal
|
||||
|
||||
Build a multi-stage AI system that learns to play *Spelunky Classic HD* by combining:
|
||||
|
||||
1. Trained perception models (custom, anchor-free CNN detectors) for entity/hazard/item recognition
|
||||
2. Hard, fast, deterministic high-level logic classifiers for HUD state (health, gold, bombs, ropes, etc.)
|
||||
3. A control agent that first imitates human input (behavior cloning) and later learns autonomously (reinforcement learning)
|
||||
|
||||
The entire perception → reasoning → action loop must fit inside a **66ms compute budget per tick** (~15 Hz), since some detection models may need to run slower than others depending on what they detect. Efficiency is a first-class design constraint, not an afterthought — every architecture decision should be made with this budget in mind.
|
||||
|
||||
A distant, non-blocking stretch goal is to eventually port the trained models to run on an FPGA (HDMI capture in, USB/game input out). This does not affect current implementation priorities but should inform model size/efficiency choices where it's free to do so.
|
||||
|
||||
---
|
||||
|
||||
## 2. System Overview
|
||||
|
||||
### 2.1 Hardware / Machines
|
||||
|
||||
| Machine | Role | Notes |
|
||||
|---|---|---|
|
||||
| New recording PC | Gameplay recording | Ubuntu 24 Desktop, dedicated to running Spelunky Classic HD fullscreen and recording sessions |
|
||||
| `jai` (existing training server) | Labeling backend, model training, inference, control agent | Ubuntu 24.04, Ryzen 7 5800X, 64GB RAM, RTX 3070 Ti (~7.45 GiB VRAM) |
|
||||
|
||||
### 2.2 High-Level Data Flow
|
||||
|
||||
```
|
||||
[Recording PC]
|
||||
Spelunky Classic HD (fullscreen, 1280x720)
|
||||
│
|
||||
▼
|
||||
Recording Script (30fps video + input log)
|
||||
│
|
||||
▼
|
||||
Lossless MP4 + input log file
|
||||
│
|
||||
▼ (transfer to jai)
|
||||
[jai — Labeling & Training]
|
||||
Web-based Labeling Tool
|
||||
├─ Set: Enemy (Main: Enemy → Sub: Bat, Snake, ...)
|
||||
├─ Set: Items
|
||||
├─ Set: Traps
|
||||
└─ Set: ... (ad-hoc, extensible)
|
||||
│
|
||||
▼
|
||||
Per-set training pipeline (anchor-free CNN detector)
|
||||
│
|
||||
▼
|
||||
Trained detection models + KPI evaluation
|
||||
│
|
||||
▼
|
||||
[jai — Runtime / Inference Loop, budget: 66ms/tick]
|
||||
Screen Capture
|
||||
│
|
||||
▼
|
||||
Parallel Detection Models (entities, per-set)
|
||||
│
|
||||
▼
|
||||
High-Level Logic Classifiers (health, gold, bombs, ropes, ... — fixed UI regions)
|
||||
│
|
||||
▼
|
||||
State Vector Builder
|
||||
(entity list + relative positions, HUD status, downsampled raw frame)
|
||||
│
|
||||
▼
|
||||
Control Agent (behavior cloning → later RL)
|
||||
│
|
||||
▼
|
||||
Input Execution (keyboard)
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 3. Component Specifications
|
||||
|
||||
### 3.1 Recording Tool
|
||||
|
||||
- **Platform:** Ubuntu 24 Desktop (new dedicated machine)
|
||||
- **Display:** Spelunky Classic HD run in fullscreen at 1280x720
|
||||
- **Input device:** Keyboard
|
||||
- **Video output:** Lossless MP4, 30 fps
|
||||
- **Input log:** Separate log file, input events mapped exactly to the 30fps video timeline (frame index + timestamp + key state)
|
||||
- **Requirements:**
|
||||
- Frame-accurate sync between video frames and logged input state
|
||||
- Output format must be easy to slice into individual frames for labeling later
|
||||
|
||||
### 3.2 Web-Based Labeling Tool
|
||||
|
||||
- **Bounding boxes only** (no segmentation/polygons needed)
|
||||
- **Hierarchical, ad-hoc class system:** Main class → Sub class (e.g. Main: `Enemy` → Sub: `Bat`, `Snake`, ...). Users can create new main/sub classes on the fly without a schema migration.
|
||||
- **Multiple independent label sets**, each targeting a specific model (e.g. `Enemy`, `Items`, `Traps`, `PositiveDetection`, ...). Frames may be shared across sets, but labels are tracked per set.
|
||||
- **Multi-user support**
|
||||
- **Simple frontend** — functionality over polish
|
||||
- **All labels must be editable** (move, resize, reclass, delete) at any time
|
||||
- **Active learning workflow:**
|
||||
1. Manually label a small seed set (e.g. ~20 instances of a class)
|
||||
2. Train an initial model on the seed set
|
||||
3. Run inference on a larger new batch (e.g. ~100 frames) to auto-label
|
||||
4. Review/correct the auto-labels in the UI
|
||||
5. Promote to a new dataset version (versioned datasets, e.g. `enemy-v1`, `enemy-v2`, ...) for reproducibility
|
||||
- **Data model must support:** sets, hierarchical classes, dataset versions, per-image/per-set label state (unlabeled / auto-labeled / reviewed)
|
||||
|
||||
### 3.3 Detection Models
|
||||
|
||||
- **Approach:** Anchor-free, center-heatmap-based detector (CenterNet-style) — NOT YOLO/anchor-based, and NOT a two-stage (R-CNN-style) detector.
|
||||
- Rationale: sprites have fairly consistent size/shape, ad-hoc new sub-classes shouldn't require anchor redesign, and heatmap peaks handle multiple same-class objects near each other well (e.g. several bats in frame).
|
||||
- Model predicts: a center-point heatmap per class + direct width/height regression from each detected center.
|
||||
- **One model per label set** (Enemy, Items, Traps, ...), trained and run independently
|
||||
- **Custom, small/efficient CNN architectures** — sized with the 66ms total-pipeline budget in mind. Multiple detection models must be able to run in parallel (or fast serial) on the RTX 3070 Ti.
|
||||
- **Inference rate:** 15 fps baseline per model, slower where acceptable depending on what's being detected
|
||||
|
||||
### 3.4 KPIs / Evaluation
|
||||
|
||||
- **Phase 1:** Standard detection metrics evaluated against held-out labeled data (precision, recall, and an mAP-equivalent adapted for the heatmap/center-point formulation)
|
||||
- **Phase 2 (later):** Ground-truth validation via direct game-code output / memory reading (approach not yet researched — open item, see Roadmap)
|
||||
|
||||
### 3.5 High-Level Logic Classifiers
|
||||
|
||||
- Covers HUD/status information: health, gold count, bomb count, rope count, level, etc.
|
||||
- Since these UI elements sit at **fixed screen positions**, use small, dedicated classifiers per fixed UI region rather than generic OCR — faster and more deterministic.
|
||||
- These feed directly into the state vector as structured status values.
|
||||
|
||||
### 3.6 Control Agent
|
||||
|
||||
- **Phase 1 — Behavior Cloning:**
|
||||
- Input (state vector) = detected entity list (class + position relative to player) + HUD status (from 3.5) + a downsampled version of the raw frame as additional context
|
||||
- Output = predicted input/action, trained to imitate the human recordings
|
||||
- **Phase 2 — Autonomous learning (later, design open):** likely reinforcement learning; reward function definition is an open item.
|
||||
- **Multi-policy design goal:** architecture should allow multiple distinct agent "personalities" to be trained later on top of the same perception stack (e.g. a "speedrunner" policy, a "score maximizer" policy), without redesigning the perception layer.
|
||||
|
||||
---
|
||||
|
||||
## 4. Timing & Performance Constraints
|
||||
|
||||
- **Hard budget: 66ms for the entire tick** — screen capture → all detection models → high-level logic classifiers → state vector construction → control agent decision → input execution.
|
||||
- Detection models should run **in parallel** wherever possible (multi-process/multi-thread or batched on GPU) to maximize how much can happen within the budget.
|
||||
- Target hardware for this budget: RTX 3070 Ti (current), assuming sufficient compute headroom; model sizes must be chosen accordingly.
|
||||
|
||||
---
|
||||
|
||||
## 5. Open Items / Roadmap
|
||||
|
||||
- [ ] Research approach for KPI validation via game memory reading (Phase 2 evaluation)
|
||||
- [ ] Define reward function(s) for the autonomous learning (RL) phase
|
||||
- [ ] Design multi-policy framework for distinct agent behaviors (speedrunner, score-farmer, etc.)
|
||||
- [ ] Decide on exact heatmap loss formulation / training details for the anchor-free detector
|
||||
- [ ] Long-term: keep model sizes/architectures FPGA-portability-aware (HDMI capture in, USB output) — not an active work item, just a soft constraint to keep in mind
|
||||
|
||||
---
|
||||
|
||||
## 6. Tech Stack & Conventions
|
||||
|
||||
- **Language:** Python (primary), for both recording tooling, labeling backend, training, and inference
|
||||
- **Environments:** Ubuntu 24 (recording PC and `jai` training server), `.venv` virtual environments
|
||||
- **All project files, code comments, and documentation in English**
|
||||
- **Documentation (READMEs) should be kept current as components are built**
|
||||
- **Git commits:** every commit message must include the name of the Claude model that authored it
|
||||
27
README.md
27
README.md
@ -0,0 +1,27 @@
|
||||
# SpelunkAI
|
||||
|
||||
A multi-stage AI system that learns to play *Spelunky Classic HD*: trained perception
|
||||
models, deterministic HUD-state classifiers, and a control agent that first imitates
|
||||
human play (behavior cloning) and later learns autonomously (reinforcement learning).
|
||||
|
||||
Full project spec, architecture, timing constraints, and roadmap live in
|
||||
[`CLAUDE.md`](./CLAUDE.md) — that document is the source of truth. This README is just
|
||||
a map of the repo.
|
||||
|
||||
## Layout
|
||||
|
||||
| Directory | Component | Spec section |
|
||||
|---|---|---|
|
||||
| [`recording/`](./recording) | Recording Tool — captures gameplay video + input log on the recording PC | 3.1 |
|
||||
| [`labeling/`](./labeling) | Web-Based Labeling Tool — backend + frontend for bounding-box labeling | 3.2 |
|
||||
| [`training/`](./training) | Per-set training pipeline for the anchor-free CNN detectors | 3.3 |
|
||||
| [`inference/`](./inference) | Runtime inference loop — capture → detectors → HUD classifiers → state vector, within the 66ms/tick budget | 3.4, 3.5 |
|
||||
| [`control/`](./control) | Control Agent — behavior cloning now, RL later | 3.6 |
|
||||
|
||||
Each component directory has its own README with more detail and its own `.venv`
|
||||
(per `CLAUDE.md` §6), since components run on different machines (recording PC vs.
|
||||
`jai`) and have independent dependencies.
|
||||
|
||||
## Status
|
||||
|
||||
Repo scaffold only — no components are implemented yet.
|
||||
19
control/README.md
Normal file
19
control/README.md
Normal file
@ -0,0 +1,19 @@
|
||||
# Control Agent
|
||||
|
||||
Consumes the state vector produced by the inference loop and predicts an action.
|
||||
Phase 1: behavior cloning, trained to imitate human input recordings. Phase 2 (later,
|
||||
design open): autonomous reinforcement learning. Architecture should allow multiple
|
||||
distinct agent "personalities" (e.g. speedrunner, score maximizer) trained later on
|
||||
top of the same perception stack.
|
||||
|
||||
See CLAUDE.md §3.6 for full requirements.
|
||||
|
||||
**Status:** not yet implemented.
|
||||
|
||||
## Setup
|
||||
|
||||
```
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate
|
||||
pip install -e .
|
||||
```
|
||||
13
control/pyproject.toml
Normal file
13
control/pyproject.toml
Normal file
@ -0,0 +1,13 @@
|
||||
[project]
|
||||
name = "spelunkai-control"
|
||||
version = "0.0.0"
|
||||
description = "SpelunkAI control agent (behavior cloning now, RL later)"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = []
|
||||
|
||||
[build-system]
|
||||
requires = ["setuptools>=68"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
where = ["src"]
|
||||
0
control/src/spelunkai_control/__init__.py
Normal file
0
control/src/spelunkai_control/__init__.py
Normal file
19
inference/README.md
Normal file
19
inference/README.md
Normal file
@ -0,0 +1,19 @@
|
||||
# Inference / Runtime Loop
|
||||
|
||||
The live tick loop: screen capture → parallel detection models (per label set) →
|
||||
high-level HUD/status logic classifiers (health, gold, bombs, ropes, level — fixed UI
|
||||
regions) → state vector construction (entity list + relative positions, HUD status,
|
||||
downsampled raw frame). Must fit the entire perception → state-vector pipeline inside
|
||||
the 66ms/tick compute budget (~15Hz).
|
||||
|
||||
See CLAUDE.md §3.4, §3.5, §4 for full requirements.
|
||||
|
||||
**Status:** not yet implemented.
|
||||
|
||||
## Setup
|
||||
|
||||
```
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate
|
||||
pip install -e .
|
||||
```
|
||||
13
inference/pyproject.toml
Normal file
13
inference/pyproject.toml
Normal file
@ -0,0 +1,13 @@
|
||||
[project]
|
||||
name = "spelunkai-inference"
|
||||
version = "0.0.0"
|
||||
description = "SpelunkAI runtime inference loop (capture, detectors, HUD classifiers, state vector)"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = []
|
||||
|
||||
[build-system]
|
||||
requires = ["setuptools>=68"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
where = ["src"]
|
||||
0
inference/src/spelunkai_inference/__init__.py
Normal file
0
inference/src/spelunkai_inference/__init__.py
Normal file
16
labeling/README.md
Normal file
16
labeling/README.md
Normal file
@ -0,0 +1,16 @@
|
||||
# Web-Based Labeling Tool
|
||||
|
||||
Bounding-box labeling tool for training frames, with a hierarchical ad-hoc class
|
||||
system (Main → Sub, e.g. `Enemy` → `Bat`), multiple independent label sets (Enemy,
|
||||
Items, Traps, ...), multi-user support, versioned datasets, and an active-learning
|
||||
workflow (seed-label → train → auto-label → review → promote).
|
||||
|
||||
See CLAUDE.md §3.2 for full requirements.
|
||||
|
||||
**Status:** not yet implemented.
|
||||
|
||||
## Structure
|
||||
|
||||
- [`backend/`](./backend) — API + data model (sets, hierarchical classes, dataset
|
||||
versions, per-image/per-set label state)
|
||||
- [`frontend/`](./frontend) — labeling UI (framework not yet decided)
|
||||
13
labeling/backend/pyproject.toml
Normal file
13
labeling/backend/pyproject.toml
Normal file
@ -0,0 +1,13 @@
|
||||
[project]
|
||||
name = "spelunkai-labeling-backend"
|
||||
version = "0.0.0"
|
||||
description = "SpelunkAI Labeling Tool backend: API + data model for bounding-box labels"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = []
|
||||
|
||||
[build-system]
|
||||
requires = ["setuptools>=68"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
where = ["src"]
|
||||
9
labeling/frontend/README.md
Normal file
9
labeling/frontend/README.md
Normal file
@ -0,0 +1,9 @@
|
||||
# Labeling Tool — Frontend
|
||||
|
||||
UI for the bounding-box labeling tool (image canvas with drawable/editable boxes,
|
||||
class picker, dataset/set navigation). Framework not yet chosen — deferred until this
|
||||
component is actively built.
|
||||
|
||||
See CLAUDE.md §3.2 for full requirements.
|
||||
|
||||
**Status:** not yet implemented.
|
||||
18
recording/README.md
Normal file
18
recording/README.md
Normal file
@ -0,0 +1,18 @@
|
||||
# Recording Tool
|
||||
|
||||
Captures *Spelunky Classic HD* gameplay on the dedicated recording PC (Ubuntu 24
|
||||
Desktop, fullscreen 1280x720) as a lossless 30fps MP4, alongside a separate input log
|
||||
mapped frame-accurately to the video timeline (frame index + timestamp + key state).
|
||||
Output must be easy to slice into individual frames for labeling.
|
||||
|
||||
See CLAUDE.md §3.1 for full requirements.
|
||||
|
||||
**Status:** not yet implemented.
|
||||
|
||||
## Setup
|
||||
|
||||
```
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate
|
||||
pip install -e .
|
||||
```
|
||||
13
recording/pyproject.toml
Normal file
13
recording/pyproject.toml
Normal file
@ -0,0 +1,13 @@
|
||||
[project]
|
||||
name = "spelunkai-recording"
|
||||
version = "0.0.0"
|
||||
description = "SpelunkAI Recording Tool: gameplay video + input log capture"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = []
|
||||
|
||||
[build-system]
|
||||
requires = ["setuptools>=68"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
where = ["src"]
|
||||
0
recording/src/spelunkai_recording/__init__.py
Normal file
0
recording/src/spelunkai_recording/__init__.py
Normal file
19
training/README.md
Normal file
19
training/README.md
Normal file
@ -0,0 +1,19 @@
|
||||
# Detection Model Training
|
||||
|
||||
Per-set training pipeline for the anchor-free, center-heatmap-based (CenterNet-style)
|
||||
CNN detectors — one model per label set (Enemy, Items, Traps, ...). Custom, small/
|
||||
efficient architectures sized around the overall 66ms/tick inference budget. Includes
|
||||
KPI evaluation (precision, recall, mAP-equivalent for the heatmap formulation) against
|
||||
held-out labeled data.
|
||||
|
||||
See CLAUDE.md §3.3–3.4 for full requirements.
|
||||
|
||||
**Status:** not yet implemented.
|
||||
|
||||
## Setup
|
||||
|
||||
```
|
||||
python -m venv .venv
|
||||
source .venv/bin/activate
|
||||
pip install -e .
|
||||
```
|
||||
13
training/pyproject.toml
Normal file
13
training/pyproject.toml
Normal file
@ -0,0 +1,13 @@
|
||||
[project]
|
||||
name = "spelunkai-training"
|
||||
version = "0.0.0"
|
||||
description = "SpelunkAI detector training pipeline (anchor-free, per-set CNNs)"
|
||||
requires-python = ">=3.10"
|
||||
dependencies = []
|
||||
|
||||
[build-system]
|
||||
requires = ["setuptools>=68"]
|
||||
build-backend = "setuptools.build_meta"
|
||||
|
||||
[tool.setuptools.packages.find]
|
||||
where = ["src"]
|
||||
0
training/src/spelunkai_training/__init__.py
Normal file
0
training/src/spelunkai_training/__init__.py
Normal file
Loading…
x
Reference in New Issue
Block a user