Discrete position/allocation policy over a small universe of liquid US equities and crypto, decided at fixed minutes-cadence intervals. Trained offline by deep Q-learning on a replay simulator built from this venue's own historical bars, with walk-forward validation only. Currently in observation phase: no orders until the sim-trained policy passes selfchecks.