Search
AI & LLM Engineering · By BagelHole
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Deploy an API gateway for LLM traffic with load balancing, rate limiting, key management, semantic caching, fallback routing, and cost tracking. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 2 | Auto-scale LLM inference clusters on Kubernetes using KEDA, custom GPU metrics, and horizontal pod autoscaling. | BagelHole/ | 1.2k | — | ~2k | Automated safety check: Pass | MIT | 4 mo ago |
| 3 | Orchestrate AI/ML pipelines for data ingestion, model training, batch inference, and RAG indexing using Prefect, Airflow, or Dagster. | BagelHole/ | 1.2k | — | ~2.3k | Automated safety check: Pass | MIT | 4 mo ago |
| 4 | Reduce LLM API and infrastructure costs through model selection, prompt caching, batching, caching, quantization, and self-hosting strategies. | BagelHole/ | 1.2k | — | ~2.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 5 | Set up infrastructure for fine-tuning LLMs with QLoRA, LoRA, and full fine-tuning using Hugging Face TRL, Axolotl, and distributed training with DeepSpeed or FSDP. | BagelHole/ | 1.2k | — | ~2.2k | Automated safety check: Pass | MIT | 4 mo ago |
| 6 | Deploy ML models on Kubernetes with KServe (formerly KFServing) and NVIDIA Triton Inference Server. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |
| 7 | Build and operate Retrieval-Augmented Generation (RAG) infrastructure with vector stores, embedding pipelines, and hybrid search. | BagelHole/ | 1.2k | — | ~2.1k | Automated safety check: Pass | MIT | 4 mo ago |