Modal Inference
PlatformModal Labsactive
Serverless and dedicated infrastructure for deploying low-latency AI model endpoints.
- —
- —
- —
- —
- —
- 2
Products
| Name | Type | Category | Launched | Latest release | Status |
|---|---|---|---|---|---|
| Product | — | Jun 2026 | Modal Auto Endpoints launched | active | |
| Product | — | Jun 2026 | Modal Servers launched | active | |
| Product | — | — | Dedicated Endpoints offered for production inference | active | |
| Product | — | — | Shared Endpoints offered for hosted models | active |
Releases
Jun 24, 2026Auto Endpoints add speculative decoding optimizationFeature update
Modal documented state-of-the-art speculative decoding performance in Auto Endpoints.
Oct 10, 2023Modal inference platform reaches GAProduct launch
Modal made serverless model deployment and inference generally available.

