Sozdai Logo
Gemini

Gemini 3.5 Flash: Online Chat & API

gemini-3.5-flash

Google's high-efficiency multimodal model — near-Pro reasoning at Flash-class speed and price, with tunable thinking levels and 1M context.

Input 255 ₽ · Output 1 530 ₽ /1M tokens1M context Thinking Prompt caching
API key
gemini-3.5-flash
Open full chat

Try it right here

Real model, streaming live — free trial credits on sign-up

Similar models

About Gemini 3.5 Flash

Gemini 3.5 Flash delivers near-Pro-level reasoning and coding at Flash-class latency and cost, and understands images natively.

Thinking depth is configurable per request — from minimal for speed to high for hard problems — and implicit caching automatically discounts repeated prefixes. The 1M-token context window handles big inputs in one call.

On Sozdai you can use Gemini 3.5 Flash two ways with one account and one credit balance: chat in the browser playground on this page, or call it programmatically through an OpenAI-compatible endpoint. Pricing is pure pay-as-you-go per token with no subscription.

Why Gemini 3.5 Flash

Tunable thinking levels

Dial reasoning from minimal to high per request — pay for depth only when the task needs it.

Fast multimodal

Native image understanding at Flash-class latency for screenshots, documents and photos.

OpenAI-compatible API

Point your existing OpenAI SDK at our base URL and it just works — streaming, tool calling, usage accounting included.

Automatic prompt caching

Repeated prompt prefixes are cached upstream automatically and billed at a fraction of the input price — no parameters needed. Great for agents and RAG.

One balance for everything

Web chat and API share the same credit balance and transparent per-token pricing. No subscription, no minimums.

Specifications

Model IDgemini-3.5-flash
FamilyGemini
TypeLLM
Context window1M tokens
Input price255 ₽ / 1M tokens
Output price1 530 ₽ / 1M tokens
Cache read price25,5 ₽ / 1M tokens
Thinking / reasoningYes
Prompt cachingYes
API compatibilityOpenAI SDK + Anthropic /v1/messages
StreamingYes

Integrate in three steps

1

Create an API key

Sign up, open the console and issue a key — free trial credits are included, no card required.

2

Point your SDK at Sozdai

Set the base URL to our endpoint and pass model \"gemini-3.5-flash\". Existing OpenAI SDK code needs no other changes.

3

Ship and monitor

Stream responses in production and track every request's tokens and cost in the usage logs.

Frequently asked questions

Can I try Gemini 3.5 Flash without writing code?+

Yes — the playground on this page is the real model. Sign up, get trial credits and chat instantly; the same account later works for the API.

Is the API compatible with the OpenAI SDK?+

Yes. Use any OpenAI SDK (Python, Node, etc.) with our base URL — streaming and tool calling included.

How does pricing work?+

Pure pay-as-you-go per token, shown on this page and billed from your credit balance (1 credit = $0.01). Cache hits are billed at a heavily discounted rate. No subscription.

Does it support thinking / reasoning?+

Yes — thinking depth is configurable per request (via reasoning effort), from quick passes to deep reasoning.

What is the context window?+

1M tokens — large inputs and long sessions in one call.