NCCL Wire Protocols
How NCCL's wire protocols trade payload bytes against synchronization cost
How NCCL's wire protocols trade payload bytes against synchronization cost
Exploring Hopper Architecture on H100 and try to match PyTorch/cuBLAS performance for GEMMs using CUTLASS
Walkthrough of optimization techniques for GEMMs from a naive fp32 kernel to CUTLASS bf16 kernel
Worklog: Performance debugging Triton Kernel
Fused Softmax Triton kernel exploration
RMS Normalization Triton kernel implementation for LLMs
Visualizer for MXFP4 quantization
KV Caching: Training vs Inference in Multi-Head Attention
Load OpenAI's reference implementation
Testing out the huggingface version for the gpt-oss-20b locally on consumer hardware