The Aether Fleet roadmap
Twelve ship systems. 35 puzzles. Bring each system online to forge its Core Shard and unlock the next, from booting the reactor to engaging the hyperdrive.
Get data onto the GPU the right way — the host↔device lifecycle the puzzles hide, and where most real-world performance is won or lost. Earns Flight Certification, not a Core Shard.
Cold Boot
Write your first GPU kernels: one thread per element, 2D grids, shared memory.
Diagnostics Bay
Hunt the bugs every GPU dev hits: bad indices, races, reads off the edge.
Systems Online
The core algorithm kit: pooling, convolution, prefix sum, tiled matmul.
Bridge Console
Package kernels as MAX custom ops, then build softmax and attention.
Fleet Standard
Ship your ops into PyTorch with a real forward and backward pass.
Drill & Clock
Make kernels fast with SIMD vectorization, then benchmark the win.
Fireteams
Reduce and shuffle data across all 32 lanes of a warp.
Deck Command
Whole block reductions, parallel histograms with atomics, normalization.
Logistics
Overlap memory and compute with async copy and pipelining.
Telemetry
Profile real kernels and tune occupancy and shared memory banking.
Hyperdrive
Drive the tensor cores and coordinate clusters of many blocks.
Final Alignment
Aligned, vectorized memory access for peak throughput.