Sozdai Logo
Gemini

Gemini 3.5 Flash Lite: Online Chat & API

gemini-3.5-flash-lite

The lightest tier of Google's Gemini line — built for massive-scale pipelines where unit cost rules, still with a 1M-token context window.

Input 51 ₽ · Output 425 ₽ /1M tokens1M context Thinking Prompt caching
API key
gemini-3.5-flash-lite
Open full chat

Try it right here

Real model, streaming live — free trial credits on sign-up

Similar models

About Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite trades peak intelligence for extreme efficiency: classification, extraction, routing, moderation and bulk chat at a price that makes millions of calls practical.

Unusually for the price it keeps a 1M-token context window and native image input, and automatic caching discounts repeated prefixes with zero configuration.

On Sozdai you can use Gemini 3.5 Flash Lite two ways with one account and one credit balance: chat in the browser playground on this page, or call it programmatically through an OpenAI-compatible endpoint. Pricing is pure pay-as-you-go per token with no subscription.

Why Gemini 3.5 Flash Lite

Built for massive scale

Unit economics that make million-call pipelines routine — the default for bulk workloads.

1M context at the bottom tier

Huge inputs and long documents remain single-call even on the lightest model.

OpenAI-compatible API

Point your existing OpenAI SDK at our base URL and it just works — streaming, tool calling, usage accounting included.

Automatic prompt caching

Repeated prompt prefixes are cached upstream automatically and billed at a fraction of the input price — no parameters needed. Great for agents and RAG.

One balance for everything

Web chat and API share the same credit balance and transparent per-token pricing. No subscription, no minimums.

Specifications

Model IDgemini-3.5-flash-lite
FamilyGemini
TypeLLM
Context window1M tokens
Input price51 ₽ / 1M tokens
Output price425 ₽ / 1M tokens
Cache read price5,1 ₽ / 1M tokens
Thinking / reasoningYes
Prompt cachingYes
API compatibilityOpenAI SDK + Anthropic /v1/messages
StreamingYes

Integrate in three steps

1

Create an API key

Sign up, open the console and issue a key — free trial credits are included, no card required.

2

Point your SDK at Sozdai

Set the base URL to our endpoint and pass model \"gemini-3.5-flash-lite\". Existing OpenAI SDK code needs no other changes.

3

Ship and monitor

Stream responses in production and track every request's tokens and cost in the usage logs.

Frequently asked questions

Can I try Gemini 3.5 Flash Lite without writing code?+

Yes — the playground on this page is the real model. Sign up, get trial credits and chat instantly; the same account later works for the API.

Is the API compatible with the OpenAI SDK?+

Yes. Use any OpenAI SDK (Python, Node, etc.) with our base URL — streaming and tool calling included.

How does pricing work?+

Pure pay-as-you-go per token, shown on this page and billed from your credit balance (1 credit = $0.01). Cache hits are billed at a heavily discounted rate. No subscription.

Does it support thinking / reasoning?+

Yes — light reasoning can be enabled per request, though the model is optimized for fast passes.

What is the context window?+

1M tokens — rare at this price tier.