Model comparison
Inkling vs NVIDIA Nemotron 3.5 Lightning
Compare Inkling and NVIDIA Nemotron 3.5 Lightning using the same provider-sourced text & reasoning rubric. No mystery score and no invented benchmark ranking.
Facts checked September 4, 2026
Model comparison
Compare Inkling and NVIDIA Nemotron 3.5 Lightning using the same provider-sourced text & reasoning rubric. No mystery score and no invented benchmark ranking.
Facts checked September 4, 2026
Set your usage. Your estimate updates as you type.
Example: 6,000 input tokens for 10 pages, plus a 500-token summary. Page lengths vary; adjust the numbers below.
One run sends your input to the model once and receives an answer. Tokens are pieces of text: input is what you send, output is the answer you receive.
Estimates exclude taxes, tools, cache storage/writes, free allowances and custom discounts. Image estimates cover output only, not prompt or reference-image charges. Quality modes differ by model. Unlisted settings are not treated as free.
| Model | Access | Estimated total (USD) |
|---|---|---|
| InklingThinking Machines Lab | Thinking Machines Lab | No reviewed rate |
| NVIDIA Nemotron 3.5 LightningNVIDIA | NVIDIA | No reviewed rate |
Quick take
Thinking Machines Lab's large open-weights model for customizable reasoning, coding, tools, vision, and audio workflows.
A 975B model is a serious serving project even with sparse activation. Compare the exact hosted or self-hosted route, and do not assume the model's 1M maximum context is available on every provider.
NVIDIA's compact 30B mixture-of-experts model for efficient specialist agents and high-volume text workflows.
The current release is labeled preview. Validate the exact precision, language, tool template, provider route, and long-context memory needs before standardizing a production fleet.
Compare the published facts
Values use each provider's own published units and limits. A blank means the provider did not publish a directly comparable value in the sources reviewed.
| Text & reasoning | Inkling | NVIDIA Nemotron 3.5 Lightning |
|---|---|---|
| Context windowMaximum combined prompt and working context documented by the provider. | Up to 1M tokens; Tinker offers 64K and 256K | Up to 1M tokens |
| Maximum outputProvider-published response limit, where available. | Not separately published | Not separately published |
| Knowledge cutoffLatest reliable knowledge date explicitly published by the model provider. Search and connected tools can retrieve newer information but do not change the model's built-in cutoff. | Not published | Pretraining through Sep 2025; post-training through May 2026 |
| Input priceCurrent standard list price per million input tokens unless noted. | Hosting dependent | Free prototype or deployment cost |
| Output priceCurrent standard list price per million output tokens unless noted. | Hosting dependent | Free prototype or deployment cost |
| InputsMedia types accepted by the listed model endpoint. | Text, image, audio | Text |
| Tools & agentsSelected native tools and agent-building capabilities, not an exhaustive list. | Agentic coding, tools, Python-assisted vision, controllable thinking | Agentic tools and long-running workflows |
How to choose
Start with the job you need to complete, then validate cost, access, and policy details on your exact provider route.
The open weights can be self-hosted so request data stays in infrastructure you control. Tinker and partner-hosted routes have separate retention and training-use terms that must be checked with the selected provider.
Self-hosted weights keep request data under the operator's controls. NVIDIA API trials, OpenRouter, and cloud partners each apply separate logging, retention, and training-use policies.
Provider and API links
Frequently asked questions
Inkling: Thinking Machines Lab's large open-weights model for customizable reasoning, coding, tools, vision, and audio workflows. NVIDIA Nemotron 3.5 Lightning: NVIDIA's compact 30B mixture-of-experts model for efficient specialist agents and high-volume text workflows.
Consider Inkling when your priority is Teams that want an adaptable open-weight foundation model. Consider NVIDIA Nemotron 3.5 Lightning when your priority is High-volume agent sub-tasks. Test both with your own data and provider route before committing.
No. This comparison aligns provider-published facts for the Text & reasoning category. It does not claim a universal winner or combine incompatible third-party benchmark scores.