Gemini 3.8 Flash: Google's fast model with built-in search
Gemini 3.8 Flash is a newer fast model in Google's Gemini 3 line. It answers in seconds and handles everyday work well: tables, code, and rewriting texts. What sets it apart from its catalog neighbors is built-in Google Search, switched on with a single toggle in the chat. It has been available on Genosai since September 19, 2026: a 120,000-token context, files and images as input, and a price of 45 / 225 ₽ per 1M tokens (input / output).
Updated: September 27, 2026
- Fast answers — A Flash-class model: responses arrive in seconds, so it is easy to keep open for a steady stream of small work tasks.
- Built-in Google Search — The Web Search toggle in the chat lets the model search Google and ground its answer in fresh data, not only in its training.
- Tables, code, and copy — Builds comparison tables reliably, writes and explains code, and rewrites or shortens texts in the tone you ask for.
- Files and images — Accepts images and files as input: screenshots, photos of documents, diagrams — and works with what is inside them.
- 120,000-token context — A long document, a conversation thread, or a large piece of code fits into one request together with the chat history.
Contents
- What is Gemini 3.8 Flash
- Capabilities
- Examples prompt and response
- How to use on Genosai
- Prompts
- Generation cost
- How it compares
- Limitations and tips
- FAQ
What is Gemini 3.8 Flash
Gemini 3.8 Flash is a newer fast model in Google's Gemini 3 line. Like other Flash-class models, it is built for speed: answers arrive in seconds rather than minutes, which makes it a handy model to keep open for a steady flow of everyday tasks. Building a comparison table, writing a small function, rewriting a text in a different tone, or cutting a long email down to one paragraph — that is typical work for Gemini 3.8 Flash.
The model arrived on Genosai on September 19, 2026. The context window is 120,000 tokens, enough for a long document, a long conversation, or a large piece of code together with the chat history. It accepts images and files as input, not just text, and responds with text. Function calling is supported.
What makes Gemini 3.8 Flash stand out among the fast models in the catalog is built-in web search. The model can search Google and ground its answer in the results. Most language models only know the world up to their training cutoff and cannot see fresh news, prices, or changes to services. Here you simply switch on the Web Search toggle in the chat, and the answer takes current information from the web into account.
The client price is 45 / 225 ₽ per 1M tokens (input / output). For comparison, its neighbor Gemini 3 Flash costs 40 / 240 ₽: the new model's input is slightly more expensive and its output slightly cheaper. It works from Russia without a VPN, with payment by Russian bank cards and no subscription.
Capabilities
Gemini 3.8 Flash is a workhorse for tasks that need speed and a clean format rather than minutes of reasoning.
Google Search right from the chat
When a question is about recent events, prices, software versions, or service terms, training data is not enough. The Web Search toggle lets the model go to Google and build its answer on what it finds. This is useful for quick market overviews, checking current terms, choosing tools, and questions whose answers change over time. Keep in mind that search requests cost more than regular ones, so turn search on where fresh data really matters.
Tables and structured answers
The model holds a requested format well. Ask for a table with specific columns and you get a table with exactly those columns, followed by exactly the conclusion you asked for. This helps with comparing options, checklists, plans, and summaries that you then paste into a document or send to colleagues.
Code
Gemini 3.8 Flash writes small functions and scripts, explains someone else's code, suggests fixes, and adds a usage example. For typical tasks — processing lists, parsing data, simple utilities — its speed and quality are usually enough on the first try. The 120,000-token context lets you pass several files and an error log at once.
Rewriting and editing
Shortening, simplifying, making the tone friendlier or more formal, rewriting for a different audience — these are tasks where fast models pay off the most. The answer comes back almost instantly, and each such request costs a fraction of a ruble.
Images and files
The model accepts images and files: an interface screenshot, a photo of a document, a diagram, or a spreadsheet. You can ask it to describe the content, find a mistake in a mockup, pull data from a scan into a table, or explain what a chart shows.
Examples prompt and response
Below are real Gemini 3.8 Flash responses obtained in the Genosai text studio, quoted without edits.
Prompt: Write a Python function that takes a list of orders like {"customer": str, "amount": float} and returns the top 3 customers by total spend. Briefly explain the solution.
Gemini 3.8 Flash response:
Here is the Python function using standard library tools (`collections.defaultdict` and `heapq`):
from collections import defaultdict
import heapq
def get_top_3_customers(orders: list[dict]) -> list[str]:
"""Calculates cumulative spend per customer and returns the top 3 customer names."""
# 1. Aggregate total spend per customer
customer_spend = defaultdict(float)
for order in orders:
customer_spend[order["customer"]] += order["amount"]
# 2. Find the 3 customers with the largest spend
top_3 = heapq.nlargest(3, customer_spend, key=customer_spend.get)
return top_3
# Example usage:
if __name__ == "__main__":
orders = [
{"customer": "Alice", "amount": 120.50},
{"customer": "Bob", "amount": 50.00},
{"customer": "Alice", "amount": 80.00},
{"customer": "Charlie", "amount": 300.00},
{"customer": "David", "amount": 150.00},
{"customer": "Bob", "amount": 25.00},
]
print(get_top_3_customers(orders))
# Output: ['Charlie', 'Alice', 'David']
Explanation
- Aggregation (`defaultdict`): Iterates through the list of orders to accumulate the total amount spent for each unique customer in $O(N)$ time.
- Top Elements Selection (`heapq.nlargest`): Efficiently retrieves the top 3 keys from the dictionary based on their aggregated values. Using a heap runs in $O(U \log k)$ time …
_Generated on Genosai.io with Gemini 3.8 Flash; the end of the answer is trimmed._
The model went straight to the right standard-library tools and added a usage example with the expected output. The output is correct: Charlie has 300, Alice 200.5, David 150, and Bob 75, so the top three are Charlie, Alice, and David.
The second example is a request in Russian: compare three ways to promote a small coffee shop (targeted ads on VK, Yandex Maps, and collaborations with local bloggers) in a table with the columns monthly budget, speed of results, and main risk, followed by one paragraph of recommendation.
Prompt: Сравни в таблице три формата продвижения для небольшой кофейни: таргетированная реклама во ВКонтакте, Яндекс Карты и коллаборации с местными блогерами. Колонки: бюджет в месяц, скорость результата, главный риск. После таблицы — один абзац с рекомендацией.
Gemini 3.8 Flash response:
| Формат продвижения | Бюджет в месяц | Скорость результата | Главный риск |
| :--- | :--- | :--- | :--- |
| **Таргетированная реклама во ВКонтакте** | 15 000 – 35 000 ₽ | Быстрая (1–3 дня с момента запуска объявлений) | Слив бюджета из-за размытого геотаргетинга или слабого оффера, не мотивирующего зайти |
| **Яндекс Карты (Геореклама)** | 10 000 – 25 000 ₽ | Средняя (2–4 недели для закрепления в топе и роста переходов) | Отсутствие визитов при низком рейтинге (ниже 4.5), плохих отзывах или скрытой вывеске |
| **Коллаборации с местными блогерами** | 5 000 – 20 000 ₽ (часто возможен бартер) | Мгновенная (всплеск в день публикации и следующие 2–3 дня) | Накрученная или нелокальная аудитория инфлюенсера, которая не конвертируется в реальных гостей |
**Рекомендация:**
Для небольшой кофейни, чей бизнес критически зависит от пешей доступности и спонтанного спроса, оптимальным стартом станет **приоритетное размещение на Яндекс Картах** — это закрывает потребность «горячей» аудитории, которая уже ищет кофе рядом. В дополнение к этому протестируйте **бартерные интеграции с местными микроинфлюенсерами** (жителями того же района), чтобы сформировать живой контент и лояльное соседское комьюнити, а таргет во ВКонтакте подключайте точечно — только под специальные события, акции выходного дня или запуск сезонного меню.
_Generated on Genosai.io with Gemini 3.8 Flash._
In short, the model estimates budgets of 15,000–35,000 ₽ for VK ads, 10,000–25,000 ₽ for Yandex Maps, and 5,000–20,000 ₽ for bloggers (often barter), and recommends starting with priority placement on Yandex Maps, adding barter deals with neighborhood micro-influencers, and using VK ads only for special events and seasonal menus. The request used 150 input and 431 output tokens and cost 0.1 ₽. Every constraint was met: three rows, the three requested columns, and a single paragraph after the table. The budget figures are the model's estimates, not market data, so check them against your own situation before deciding.
How to use on Genosai
- Open the Genosai studio on the Text tab and sign in.
- Pick Gemini 3.8 Flash in the model list.
- If needed, open the chat settings panel: temperature, top_p, a response limit of up to 32,768 tokens, and a system prompt.
- If you need an answer based on fresh data, switch on the Web Search toggle.
- Describe the task: goal, context, and answer format. Attach an image or file if it helps.
- Read the answer and refine it in the same thread — the history is kept within the 120,000-token context.
Billing is per token with no subscription, and you can top up your balance with a Russian bank card without a VPN.
The settings are simple. Gemini 3.8 Flash has no separate reasoning effort option — the model always answers in fast mode. Temperature controls variety: keep it lower for tables, code, and facts, and raise it for creative copy variations. A system prompt is handy when you need the same format many times in a row: set the role and answer structure once, then change only the input data.
Gemini 3.8 Flash is available only in the studio on the website. It is not in the Genosai public API yet, so for integrations into your own services, pick another model from the catalog.
Prompts
Templates that play to the model's strengths: speed, format, and search. Replace the values in angle brackets with your own data.
Compare <options> in a table with the columns: price, timeline, main risk. After the table, give one recommendation in a single paragraph.
Rewrite this text shorter and simpler, keeping all the facts. Friendly tone, no jargon: <text>
Write a Python function that <task>. Add a usage example and briefly explain the solution.
With web search: find the current terms of <service or plan> and list them with the date they apply from.
Look at this screenshot and list what prevents a user from completing the form.
Turn these notes into a weekly checklist with priorities: <notes>
Generation cost
The client price of Gemini 3.8 Flash on Genosai is 45 / 225 ₽ per 1M tokens (input / output). This is a promotional price from the provider, valid until December 31, 2026; after that date, prices may change. The current cost is always shown in the studio before you send a request.
Requests with web search turned on cost more than regular ones: the pages found online are added to the request and billed as input tokens. So it makes sense to enable search for questions where fresh data matters and keep it off for editing, code, and work with text you have already provided.
For a rough budget: the coffee-shop table example above — 150 input and 431 output tokens — cost 0.1 ₽. A request with 2,000 input tokens and a 1,000-token answer costs about 0.3 ₽, or about 0.5 ₽ with search. There is no subscription; only actual usage is charged. Current prices are in the Pricing section.
How it compares
Gemini 3.8 Flash is a fast, mid-priced model with built-in search. Here is how it stacks up against the closest models in the catalog.
| Model | Price ₽ per 1M (input / output) | Context | Niche |
|---|---|---|---|
| Gemini 3.8 Flash | 45 / 225 | 120,000 | Fast answers and Google Search |
| Gemini 3 Flash | 40 / 240 | 120,000 | Fast reasoning, available in the API |
| Gemini 3.1 Flash Lite | 20 / 120 | 120,000 | Cheapest in the line |
| Gemini 3.1 Pro | 160 / 960 | 120,000 | Deep analysis |
| GPT-6 Luna | 8 / 40 | 256,000 | High-volume tasks |
| DeepSeek V4 Flash | 1.12 / 52.8 | 800,000 | Very large volumes of text |
| Claude Haiku 4.5 | 80 / 400 | 180,000 | Anthropic's fast model |
Gemini 3 Flash is the closest neighbor: its input is 12.5% cheaper (40 vs 45 ₽), while its output is about 6% more expensive (240 vs 225 ₽), so Gemini 3.8 Flash comes out slightly cheaper on tasks with long answers. Gemini 3 Flash is also available in the public API, while Gemini 3.8 Flash is studio-only. Gemini 3.1 Flash Lite is 2.25 times cheaper on input and almost half the price on output — pick it for simple bulk tasks. Gemini 3.1 Pro costs about 3.6 times more on input and 4.3 times more on output, but is better suited for complex analysis.
If the lowest price is what matters, look at GPT-6 Luna: it is about 5.6 times cheaper on both input and output and has a 256,000-token context. DeepSeek V4 Flash is almost 40 times cheaper on input and 4.3 times cheaper on output, with an 800,000-token context for very long documents. Claude Haiku 4.5 costs about 1.8 times more than Gemini 3.8 Flash. The full list of models is in the model catalog.
Limitations and tips
Gemini 3.8 Flash is a fast model and has the limitations typical of its class. For complex multi-step reasoning, tangled logic, and large research tasks, Gemini 3.1 Pro is a better fit. Gemini 3.8 Flash has no separate reasoning mode, so for tricky problems ask the model to work through the solution step by step right in the prompt.
As with any fast model, the provider can be temporarily unavailable at peak times. If the studio shows a temporarily unavailable message, retry in a minute or switch to Gemini 3 Flash, which is close in speed and price.
Web search broadens what the model knows, but it does not make the answer error-free: the pages it finds can also be outdated or inaccurate. For important facts, ask the model what it relied on and check key figures against the original source. Without search, the model does not know fresh news, exchange rates, or prices — pass such data in the request yourself.
State the answer format explicitly: a table with named columns, a number of points, a length limit. The model holds such constraints well. For code, pass the full context — related files, library versions, the error text.
Most importantly, verify the result manually, especially when money, security, or important decisions depend on it. Run the code and check numbers and facts — even a careful model sometimes makes mistakes.
FAQ
What is Gemini 3.8 Flash?
It is a newer fast model in Google's Gemini 3 line, built for everyday tasks where speed matters: tables, code, rewriting texts, and answering questions. It has been available on Genosai since September 19, 2026, with a 120,000-token context and built-in Google Search.
How much does Gemini 3.8 Flash cost on Genosai?
The client price is 45 / 225 ₽ per 1M tokens (input / output). This is a promotional price from the provider, valid until December 31, 2026; after that, prices may change. Requests with web search cost more: the pages found online are added to the request and billed as input tokens. There is no subscription: you pay only for the tokens you use.
How do I turn on web search?
Open the text studio, pick Gemini 3.8 Flash, and switch on the Web Search toggle in the chat. The model can then search Google and base its answer on what it finds. With the toggle off, it answers only from its own knowledge and what you pass in the request.
How does Gemini 3.8 Flash differ from Gemini 3 Flash?
Both are fast models with a 120,000-token context. Gemini 3.8 Flash is the newer model in the line and has built-in Google Search. On price, its input is slightly more expensive (45 vs 40 ₽ per 1M) and its output is cheaper (225 vs 240 ₽ per 1M).
Can I use Gemini 3.8 Flash through the API?
No, this model is not available in the Genosai public API yet — you can use it only in the text studio on the website. If you need an API integration, pick another model from the catalog, such as Gemini 3 Flash.
Does the model have a reasoning effort setting?
No, Gemini 3.8 Flash has no such setting. The available options are temperature, top_p, a response limit of up to 32,768 tokens, and a system prompt. The model always answers in fast mode, without a separate long-thinking mode.