Repositories list
68 repositories
Hyperloom
PublicPrimus-Turbo
PublicA high-performance acceleration library dedicated to large-scale model training on AMD GPUsInfera
PublicMore token goodput from frontier models. A distributed, SLA-aware serving mesh — disaggregated prefill/decode, KV-aware routing, and cache offload, tuned to you…GEAK
PublicPrimus-SaFE
PublicPrimus-SaFE(Stability and Fault Endurance)TraceLens
PublicAutomating analysis from trace filesPrimus
PublicA flexible and high-performance training framework designed for large-scale foundation model training on AMD GPUs- Diffusion model inference benchmarking, profiling and optimizing
AgentKernelArena
PublicAgentKernelArena provides an end-to-end siloed-benchmarking environment where different LLM-powered agents—such as Cursor Agent, Claude Code, Codex, SWE-agent, …Apex
PublicAgents, and RL environment, for optimizing GPU kernels on AMD ROCm using LLM agents. Benchmarks LLM serving workloads end-to-end, profiles bottleneck kernels, o…ALTO
PublicALTO: Advanced Low-precision Training and OptimizationMagpie
PublicA lightweight, general-purpose framework for evaluating GPU kernel and benchmark.vllm-2026
PublicInstella-MoE
PublicAMDLongContextServing
PublicKimi-Linear long-context FP8 benchmark tooling for vLLM on AMD MI355X (AITER MLA head-padding + patched decode/prefill kernels).pr_pundit
PublicAn agent for OSS contributionsmlperf-common
PublicFarSkip-Collective
PublicTraining and inference implementation of FarSkip-Collective models enabling communication-computation overlapmaxtext-slurm
PublicHummingbirdXT
PublicThis repository presents an efficient acceleration pipeline for Diffusion Transformer (DiT) based video generation models, optimized for AMD client-grade GPUs, …torchtitan-amd
PublicA PyTorch native platform for training generative AI modelsgpt-fast
PublicInstella-Math
PublicAMD-LLM
PublicTraining code and resources for AMD-135M language models on AMD GPUs.m3d_rocm
PublicThis project is an optimized version of Matrix3D. It has better compatibility with ROCm ecosystem.prime_amd
PublicPARD
PublicPARD: Accelerating LLM Inference with Low-Cost PARallel Draft Model Adaptation (ICLR 26)ReasonLite
PublicEfficient Reasoning Modelssd
PublicA lightweight inference engine supporting speculative speculative decoding (SSD).DynamicChunkingDiT
PublicOfficial code for DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking
ProTip! When viewing an organization's repositories, you can use the
props. filter to filter by custom property.