Native minimax-h3 inference for apple silicon. specialized 256-thread kernel gathers and quantizes each h3 row directly into the projection's row-major int8 buffer, eliminating the intervening full-width bf16 transpose without changing any output byte. cross-m5 measurements improved complete 512 forwards by about 0.2-0.8% avoiding a second device-memory read when