Add ffmpeg/x11grab-based video capture (libx264rgb, qp=0, true lossless RGB) and an evdev-based keyboard state logger, orchestrated by a single frame-tick loop so each JSONL input row lines up 1:1 with its video frame. Unit-tested with fake keyboard/video components (no real device or ffmpeg needed); real hardware capture still needs validation on the recording PC. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
SpelunkAI
A multi-stage AI system that learns to play Spelunky Classic HD: trained perception models, deterministic HUD-state classifiers, and a control agent that first imitates human play (behavior cloning) and later learns autonomously (reinforcement learning).
Full project spec, architecture, timing constraints, and roadmap live in
CLAUDE.md — that document is the source of truth. This README is just
a map of the repo.
Layout
| Directory | Component | Spec section |
|---|---|---|
recording/ |
Recording Tool — captures gameplay video + input log on the recording PC | 3.1 |
labeling/ |
Web-Based Labeling Tool — backend + frontend for bounding-box labeling | 3.2 |
training/ |
Per-set training pipeline for the anchor-free CNN detectors | 3.3 |
inference/ |
Runtime inference loop — capture → detectors → HUD classifiers → state vector, within the 66ms/tick budget | 3.4, 3.5 |
control/ |
Control Agent — behavior cloning now, RL later | 3.6 |
Each component directory has its own README with more detail and its own .venv
(per CLAUDE.md §6), since components run on different machines (recording PC vs.
jai) and have independent dependencies.
Status
Repo scaffold only — no components are implemented yet.
Description
Languages
Python
83.1%
JavaScript
12.7%
CSS
2.3%
HTML
1.9%