Documentation
Antfly Inference
Local ML inference for ONNX-based models with an Ollama-compatible API — embeddings, chunking, reranking, named entity recognition, and text rewriting. Runs standalone or inside an Antfly cluster.
Getting Started
Install Antfly Inference, pull models, and serve your first requests
API Reference
Complete REST API documentation for the inference endpoints
Embedding Models
Generate text and multimodal vector embeddings
Reranking
Re-score search results for relevance
Chunking
Split text into semantically coherent segments
Kubernetes Operator
Deploy Antfly Inference on Kubernetes with autoscaling
Models
Browse available embedding, reranking, and chunking models
Downloads
Download Antfly Inference for your platform