Together AI · First-party API · LLM

thinkingmachines/Inkling

API availabilityAvailable in the API since Jul 16, 2026

Notice

API availability

Available in the API since Jul 16, 2026

Together AI announced API availability for this model on 2026-07-16. serverless

This date applies to API availability. Rollouts in consumer products or other hosting routes can happen on different dates.

Hosting route
First-party API
Affected scope
serverless
Announced
Jul 16, 2026
First seen by ModelClock
Oct 7, 2026

What the provider published

docs.together.ai ↗

​ New serverless models

The following models are now available on serverless :

  • thinkingmachines/Inkling: 552.8B parameters, 524,288 context length, NVFP4 quantization. Pricing: $1.00 input / $4.05 output / $0.17 cached input (per 1M tokens).

Read from the provider's text; the quote is the provider's exact lines.

Revision history

Loading…
Raw history JSON