Ollama is a local LLM runtime that lets you pull, run, and customize models, offering a CLI and REST API for chat, generation, and embeddings.

179.5k stars17.6k forksMITlast commit Actively maintained
Cloud API providing access to OpenAI models (text, chat, embeddings, images, audio) for inference, content generation, embeddings extraction, image/audio generation and fine-tuning; used to build assistants, RAG systems, and automation in applications.
Ollama is a local LLM runtime that lets you pull, run, and customize models, offering a CLI and REST API for chat, generation, and embeddings.

179.5k stars17.6k forksMITlast commit Actively maintained
Run LLMs, image, and audio models locally with an OpenAI-compatible API, optional GPU acceleration, and a built-in web UI for managing and testing models.

48.7k stars4.4k forksMITlast commit Actively maintained
Open-source Python framework to build, scale, and deploy multimodal AI services and pipelines with gRPC/HTTP/WebSocket support and Kubernetes/Docker integration.

21.9k stars2.2k forksApache-2.0last commit Slowing down
Self-hosted, OpenAI API-compatible server for streaming transcription, translation, and speech generation using faster-whisper and TTS engines like Piper and Kokoro.

3.6k stars444 forksMITlast commit Actively maintained
Every option on this page is open source and free to run on your own hardware, so you own the data and there is no subscription to cancel. 3 of 4 shipped a commit in the last six months. Licences in this list: MIT, Apache-2.0. In exchange you take on hosting, backups and updates yourself.
Browse everything in Model Serving & Inference.