We build the measurement apparatus before the strategy. Order-book depth at ten levels across four venues, eleven years of recorded ticks, stored raw and replayed in the order it arrived.
Tick-level reconstruction with venue-accurate queue position and latency. A quote placed in replay is filled the way it would have been filled, or it is not filled at all.
Sim-to-real validation on held-out periods the researcher never had access to. Divergence between simulated and realised fills is logged as an error in the environment, not in the market.
Graded research problems with a fixed scoring rule and a sealed test split. The grade is produced by the environment; nobody negotiates it afterwards.