jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Python20K stars1.7K forks+60 today
Topics
#apple-silicon#inference-server#llm#macos#mlx#openai-api
What it does
oMLX is a Python-based LLM inference server optimized for Apple Silicon, featuring continuous batching and tiered caching. It offers convenience and control through macOS menu bar management.
Star history
Not enough history yet — 1 day(s) recorded. The daily snapshot builds this up.
Tracking
- Last trending
- 2026-08-17
Creator kit
Hook
🚀 Discover oMLX, the ultimate LLM inference server for Mac users who want seamless integration and control over their AI models. #oMLX #AI
Content angles
- Create a tutorial on setting up oMLX with your favorite LLM model.
- Compare performance benchmarks of oMLX against other LLM servers on Apple Silicon devices.
- Develop a guide for integrating oMLX into coding workflows, enhancing productivity with AI assistance.
Who should care
AI enthusiasts, developers, and data scientists working on macOS who need efficient local LLM inference.