Andyyyy64/whichllm
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Explore 22 GitHub repositories focused on gpu. Discover top-starred projects and those trending this week.
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
FlashInfer: Kernel Library for LLM Serving
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Linux GPU Configuration And Monitoring Tool
Optimized primitives for collective multi-GPU communication
Achieve state of the art inference performance with modern accelerators on Kubernetes
Up to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops 🦖
Beautiful, open source, WebGPU-based charting library
Interactive data visualizations and plotting in Julia
Tensors and Dynamic neural networks in Python with strong GPU acceleration
Bend 2: a fast language that blocks AI mistakes via proof. Install: curl -fsSL https://bend-lang.com/install.sh | sh
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
FlashInfer: Kernel Library for LLM Serving
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Linux GPU Configuration And Monitoring Tool
Optimized primitives for collective multi-GPU communication
Achieve state of the art inference performance with modern accelerators on Kubernetes
Up to 100x faster strings for C, C++, CUDA, Python, Rust, Swift, JS, & Go, leveraging NEON, AVX2, AVX-512, SVE, GPGPU, & SWAR to accelerate search, hashing, sorting, edit distances, sketches, and memory ops 🦖
Beautiful, open source, WebGPU-based charting library
Interactive data visualizations and plotting in Julia
CUDA Core Compute Libraries
LUPINE is a GPU over IP bridge allowing GPUs on remote machines to be attached to CPU-only machines.
Distributed AI Model Training and LLM Fine-Tuning on Kubernetes
Open Source Inference Research Platform Standard / 开源推理研究平台
Estimate whether a Hugging Face model fits and fine-tunes on your local GPU.
Use your NVIDIA GPU's VRAM as swap space on Linux. Built for laptops with soldered memory and no upgrade path. If you have an RTX card sitting there with 8GB of VRAM and you're getting swapped to SSD, this puts that VRAM to work
A minimal, cross-compatible CPU/GPU telemetry monitor with accurate data directly from vendor APIs and beautiful ASCII visualization.
The native macOS terminal that keeps your sessions running and tells you when a coding agent needs you. GPU-rendered, scriptable, agent-aware.
Project Titania is a complete large language model system, from transformer to transistor, simple enough for one person to understand.