Skip to content
AIMarketCap

Together Inference Engine

Product

Optimized inference runtime powering Together endpoints with research-driven kernels, quantization, and speculative decoding.

Launched
—
Models
—
Modality
—
Open source
—
API available
—
Releases
2

Releases

2
Jul 18, 2024Together releases Inference Engine 2.0Together Inference Engine · 2.0

Together released a new inference stack with Turbo and Lite endpoints and reported decoding throughput up to four times faster than vLLM.

Nov 2023Together launches Inference EngineTogether Inference Engine

Together introduced its optimized inference engine across serverless endpoints and dedicated model deployments.