FeynmanWiki
ExploreLibraryRoadmapsCreateMy blogsPricing

Library

Search the public knowledge base.

Filters

Reset all

Categories

Machine Learning22Reinforcement Learning11Transformers5LLMs4Attention3Diffusion2AI Interpretability1BPE1causal interventions1context paradigm1Cordis1Databases1Distributed Machine Learning Systems1Distributed Training1DTensor1Dynamic software composition1energy-based models1FSDP1GPU-Parallel Robot Learning1importance sampling1

Sort by

LatestMost readRecently updated

Library

Search the public knowledge base.

K
Machine LearningReinforcement LearningTransformersLLMsAttentionDiffusionAI InterpretabilityBPEcausal interventionscontext paradigm

1 results

Sort by
LatestMost readRecently updated

veScale-FSDP: Flexible, Structure-Aware Sharding at 10K-GPU Scale

Suppose you want to train a model with a matrix optimizer such as Muon. The optimizer does not think of a weight matrix as a bag of ...

Distributed Machine Learning SystemsFSDPRaggedShard
Sep 2, 2026 20 min read 125