vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Explore 11 GitHub repositories focused on transformer. Discover top-starred projects and those trending this week.
A high-throughput and memory-efficient inference and serving engine for LLMs
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Fully automatic censorship removal for language models
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Production First and Production Ready End-to-End Speech Recognition Toolkit
Scalable and user friendly neural :brain: forecasting algorithms.
基于深度强化学习的开源自动因子工厂。
MoBA: Mixture of Block Attention for Long-Context LLMs
Convert images of LaTex math equations into LaTex code.
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
A high-throughput and memory-efficient inference and serving engine for LLMs
Fully automatic censorship removal for language models
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
Production First and Production Ready End-to-End Speech Recognition Toolkit
Scalable and user friendly neural :brain: forecasting algorithms.
基于深度强化学习的开源自动因子工厂。
MoBA: Mixture of Block Attention for Long-Context LLMs
Convert images of LaTex math equations into LaTex code.
[ICML 2024] LLMCompiler: An LLM Compiler for Parallel Function Calling
Tiny Model, Big Logic: Diversity-Driven Optimization Elicits Large-Model Reasoning Ability in VibeThinker-1.5B