debpalash/VoiceStudio
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Explore 21 GitHub repositories focused on cuda. Discover top-starred projects and those trending this week.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
A high-throughput and memory-efficient inference and serving engine for LLMs
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
The open-source AI voice studio. Clone, dictate, create.
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
FlashInfer: Kernel Library for LLM Serving
Optimized primitives for collective multi-GPU communication
Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
LLM speculative inference server for consumer & heterogeneous hardware
CUDA Core Compute Libraries
LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.
A high-throughput and memory-efficient inference and serving engine for LLMs
The open-source AI voice studio. Clone, dictate, create.
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
FlashInfer: Kernel Library for LLM Serving
Optimized primitives for collective multi-GPU communication
Train, inspect, edit, automate, and export 3D Gaussian Splatting scenes from a single native application.
Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end.
LLM speculative inference server for consumer & heterogeneous hardware
CUDA Core Compute Libraries
LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.
Open Source Inference Research Platform Standard / 开源推理研究平台
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
Write Java. Run on GPUs. Fast.
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
Real-time 3D full-body reconstruction from a single camera, Multiperson BVH output, Pure C++ runtime, ONNX + ggml, 70-joint skeleton with hands.
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
Use your NVIDIA GPU's VRAM as swap space on Linux. Built for laptops with soldered memory and no upgrade path. If you have an RTX card sitting there with 8GB of VRAM and you're getting swapped to SSD, this puts that VRAM to work
Local AI song generator with an editable score — YuE2 on your GPU: full songs with vocals, sheet music, covers, exact replay. Native Windows app, no Python, installer with auto-update.