oMLX is a native macOS inference server leveraging MLX for optimized LLM performance on Apple Silicon. It features paged SSD KV caching and continuous batching, significantly reducing agent response times from minutes to seconds. It offers an OpenAI and Anthropic compatible API, making it a drop-in solution for various AI tools and workflows.
