Serverless Inference Invoice from Digital Ocean (2026-07)

in #blog5 days ago

Total Usage Charges: $0.51 from Digital Ocean Serverless Inference

image.png

I recently experimented with several large language models through a serverless inference platform, including DeepSeek R1 Distill Llama 70B, Llama 3.3 70B, Qwen 3.5, Qwen3-32B, Kimi K2.5/K2.6 and MiniMax M2.5. Despite trying multiple models and processing approximately 382,000 tokens, the total bill came to only $0.51. This shows just how affordable serverless AI inference can be when there is no need to reserve or maintain dedicated GPU infrastructure.

Most of the expense came from output tokens rather than input tokens. The largest individual charge was $0.20 for generating 100,302 tokens with Kimi K2.5. By comparison, MiniMax M2.5 produced a similar 100,933 output tokens for only $0.09, mainly because its off-peak output rate was considerably lower. Input-token charges were almost negligible, with most individual items rounded to just one cent.

The invoice also highlights how much model pricing can vary, even when the workloads are broadly similar. Choosing a model is therefore not only about intelligence, speed and output quality; token pricing and off-peak discounts can make a substantial difference at scale. For small experiments, the difference may only be a few cents, but for applications generating millions or billions of tokens, selecting the right model and running workloads during off-peak periods could lead to significant savings.

Steem to the Moon🚀!

Support me, thank you!

Why you should vote me? My contributions
Please vote me as a witness or set me as a proxy via https://steemitwallet.com/~witnesses

image.png

Sort:  

Your experiment highlights the flexibility and cost-effectiveness of serverless inference, especially when combined with the right model. I'm curious, how do you plan to apply these findings to future AI model testing? 💻🤖