Search
DeepSeek · Model hubs and datasets
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Adapt and port new LLM model architectures to this xinfer project. | guoqingbao/ | 334 | — | ~4.2k | Automated safety check: Notes | MIT | 1 mo ago |
| 2 | Draw Sebastian-Raschka-gallery-style TikZ architecture diagrams for any HuggingFace decoder-only LLM, with per-block parameter formulas and concrete numbers. | yzlnew/ | 149 | — | ~2k | Automated safety check: Pass | No licence | 3 mo ago |
| 3 | Estimate GPU memory usage for Megatron-based MoE (Mixture of Experts) and dense models. | yzlnew/ | 149 | — | ~2.2k | Automated safety check: Pass | No licence | 3 mo ago |
| 4 | Train Mixture of Experts (MoE) models using DeepSpeed or HuggingFace. | Orchestra-Research/ | 13k | 2 repos | ~3.7k | Automated safety check: Pass | MIT | 3 mo ago |
| 5 | Set up the minimal set of artifacts (tokenized DCLM corpus shard + released HuggingFace checkpoint converted to DCP) required to benchmark, profile, or regression-test a MoE model in PithTrain. | mlc-ai/ | 355 | — | ~399 | Automated safety check: Pass | Apache-2.0 | 2 days ago |
| 6 | A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping. | ericrisco/ | 180 | — | ~4.1k | Automated safety check: Pass | MIT | yesterday |