Majestic Labs Software Development Planning

HW & Emulations Delivery Firmware System & Core AI Kernels Compiler & Toolchain Model Zoo
6
HW & Emulations Delivery · 0 done
1
Firmware · 0 done
30
System & Core · 0 done
21
AI Kernels · 0 done
40
Compiler & Toolchain · 0 done
1
Model Zoo · 0 done
99
Total items
Vertical / MonthJun 26v0.5
Jul 26v0.6
Aug 26v0.7
Sep 26v0.8
Oct 26v0.9
Nov 26v0.95
Dec 26v1.0
Jan 27v1.1
Feb 27v1.2
Mar 27v1.3
Apr 27v1.4
May 27v1.5
Jun 27v1.6
Jul 27v1.7
Aug 27v1.8
Sep 27v1.9
Oct 27v2.0
Nov 27v2.1
Dec 27v2.3
HW & Emulations Delivery
+ add
Palladium
+ add
+ add
+ add+ add+ add
Tape Out
+ add
+ add+ add+ add+ add+ add
Evaluation Boards
+ add
+ add
Full Majestic System in our lab
+ add
+ add+ add+ add
Customer Early Access
+ add
Firmware
SCP boot flow closed
+ add
+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add
System & Core
E2: Core Graph Execution
+ add
E2: XDMA Engine
+ add
+ add
+ add
E2: ARM QEMU Env
MDD: API ready
+ add
E2: Kernel driver for device access
MDD: Initial features (?)
+ add
MDD: User Space mgmt Tool
+ add
+ add
+ add
MDD: MMU (device side)
+ add
E2: Graph Optimizations
+ add
E2: Fault tolerance
+ add
+ add
E2: Data digestion support
+ add
E2: Advanced Scheduling
+ add
E2: Multi-Chip / Multi-E2
+ add
+ add+ add+ add
AI Kernels
+ add
Func: GEMM BF16; Softmax
Perf: GEMM FP32; RMSNorm
Precision: exp; reciprocal
+ add
Func: FlashAttention; GEMM FP8, FP4 with Blocked Scaling
Perf: GEMM BF16, FP8; Softmax
SystemC: Performance Regression Suite
+ add
+ add
+ add
+ add
+ add
P0.0 perf: rng; pointwise; Reduce; RoPE
Conv3D; Conv1D (dilated); ConvTranspose1D; interpolate
+ add
+ add
+ add
+ add
+ add
+ add
+ add
+ add
+ add
+ add
+ add
+ add
Compiler & Toolchain
Triton: Raise grid / scratchpad caps
+ add
+ add
+ add
+ add
+ add
+ add
+ add
+ add
Triton: Sliding-Window Attention (window=512)
Triton: FP8 KV-cache load/store kernels
+ add
Triton: MoE top-8 routing + expert dispatch
+ add
+ add
Triton: Flash-decode 131k context
+ add
+ add+ add+ add+ add+ add+ add+ add
Model Zoo
Laguna - Functional
+ add
+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add+ add
SystemC MLAMP→ Aug 26
CL: Simple host to ML1 communication layer, replace current TCP implementation→ Sep 26
MOS: Buffer Allocation & Kernel Metadata→ Sep 26
MDD: Planning→ Sep 26
E2: Buffer Allocation & Kernel Metadata→ Aug 26
E2: SOA, Boot & Core Graph Execution→ Sep 26
CL: Finalize host to ML1 communication layer→ Dec 26
E2: Advanced Scheduling (1)→ Dec 26
E2: Advanced Scheduling (2 - NUMA)→ Feb 27
CL: Multiple ML1 connection, trivial orchestration (round robin)→ Mar 27
MOS: MMU (device side)→ Mar 27
MDD: Graph Execution Feature Complete→ Feb 27
E2: MMU (device side)→ Mar 27
CL: Multiple ML1 orchestration (?)→ Jun 27
MOS: Fault tolerance (?)→ Jun 27
MDD: Fault Tolerance→ Jun 27
E2: Debugging→ Jul 27
MDD: Data Digestion Support→ Sep 27
Interface Definition for Graph Compiler→ Jul 26
GEMM Fused-Epilogue (bias + act); Activation; Pointwise (Add/Op/Scale/Set); Reduce; LayerNorm→ Oct 26
Fused rms_norm; layer_norm + residual; RoPE (rotary); fused_dense / MLP→ Oct 26
mha_varlen_fwd; mha_fwd_kvcache (paged-KV decode); FA2 feature axes; SDPA_attributes→ Oct 26
rng: Philox core, state mgmt, uniform draw, normal draw→ Oct 26
P0.0 perf: attention (mha_fwd / varlen / kvcache); GEMM fused-epilogue→ Jan 27
Conv2D forward (3×3, 1×1); GroupNorm→ Jan 27
Kernel Infra Performance→ Mar 27
Pooling; BatchNorm; FFT (1D complex, 1D real, STFT)→ Mar 27
P0.0 Performance Cleanup→ Mar 27
P1 perf: Conv2D / 3D / 1D; interpolate; FFT→ May 27
P1 perf: Pooling, BatchNorm, GroupNorm, ConvTranspose1D; final P0.0 + P1 performance pass→ May 27
Robustness / Integration / Stress Testing→ Dec 27
FE: Slicing Algorithm→ Aug 26
FE: Cluster Algorithm→ Sep 26
FE: DMA Operator Support→ Jul 26
FE: MCMA Support→ Oct 26
FE: factory-op rewrite productionised→ Jul 26
FE: MCMA wrapper as Inductor codegen target→ Aug 26
Triton: Adopt TritonGPU dialect→ Jul 26
Triton: Port majestic codegen onto TritonGPU IR→ Aug 26
Triton: Link riscv64 libTritonMajesticRuntime→ Jul 26
Triton: Atomic-op codegen to XM via GW→ Jul 26
FE: Cluster Algorithm→ Sep 26
BE: FX → MLIR Translation→ Jul 26
BE: Cluster→ Aug 26
BE: MajesticRunner→ Jul 26
BE: MCMA Lowering & Generation→ Sep 26
FE: Inductor auto-generated code for Majestic→ Nov 26
FE: ATen native dispatch + device register→ Sep 26
BE: XDMA Function Support (in / out kernel)→ Sep 26
BE: Memory Materialisation & Bufferisation→ Oct 26
FE: ATen factory + elementwise→ Oct 26
Triton: Tile / warp / num_stages → CMA shape mapping→ Sep 26
Triton: Validate 88-op set vs TritonCPU baseline→ Sep 26
Triton: Tile-size sweeps for top-8 hot ops→ Oct 26
Triton: FP8 / BF16 / INT4 dtype lowering→ Oct 26
BE: Kernel Lowering — AOT→ Nov 26
Triton: Decommission TritonCPU path→ Nov 26
BE: Kernel Lowering — JIT (EW · Matmul · Reduction)→ Feb 27
FE: ATen Linear/SDPA hand-tuned→ Feb 27
BE: LLVM Backend Optimisation (RISC-V codegen)→ Mar 27
FE: PTMage cost-model on FXIR→ Feb 27
Triton: Block-pointer + swizzling in GEMM→ Feb 27
Triton: Triton → Cortex MLIR with CMA hints→ Mar 27
Triton: Persistent-kernel rewrite (Linear/SDPA)→ Feb 27
BE: Full CMA Catalog + PTMage Cost Model→ Feb 27
Triton: SW-pipeline GM→XM xDMA with compute→ Apr 27

Appendix A — AI Kernels: Full Operator Classification by Library

Source: kernel-roadmap §2 · Priority 0 = LLM must-have → 5 = classical sci-compute · — = not supported

Summary — Operator counts by library and priority

LibraryP0P1P2P3P4P5Not supportedTotal
blas — cuBLAS / rocBLAS70000264073
dnn — cuDNN / MIOpen690060425
flash-attention — FlashAttention v2/v3 + cuDNN fMHA1700030121
thrust-cub — Thrust + CUB0034000135
sparse — cuSPARSE / rocSPARSE00014001024
graph-analytics — cuGraph + cugraph-ops / pylibcugraphops0006300366
linalg-solvers — cuSOLVER / hipSOLVER00000273966
rng — cuRAND / hipRAND1500004928
fft — cuFFT / rocFFT0120003520
dataframes — cuDF0006700168

Priority legend: P0 LLM inference must-have   P1 Diffusion / video / audio   P2 Thrust/CUB parallel primitives   P3 GNN / tabular NN   P4 Training-only   P5 Classical sci-compute   No plan to support