Nano Banana 2.1 is Google's fast image model for edits and posters. On the paid API a 1K picture is $0.0336, and there is no free tier.
The picture got cheaper. The words did not.
Google's Gemini API pricing page, read on 8 October 2026, prices gemini-nano-banana-2.1 with no free tier on standard or batch. Paid rates are per million tokens. The picture on this listing is the share image from the DeepMind product page, including a hydrangea sample that page pairs with a long prompt. It is not a result from this site.
- Standard input: $1.50 for text, image, or video. Text and thinking output: $7.50. Image output: $30.
- A 1K image, 1024 pixels, is 1,120 tokens, which they print as $0.0336. A 2K image, 2048 pixels, is 1,680 tokens, $0.0504. A 4K image, 4096 pixels, is 3,780 tokens, $0.113.
- Batch is half: input $0.75, text and thinking $3.75, images $15. That is $0.0168, $0.0252, and $0.0567 at 1K, 2K, and 4K.
- The same page's Nano Banana 2 row, model
gemini-3.1-flash-image, is input $0.50, text and thinking $3, and images $60. Its sizes are $0.045 at 0.5K (747 tokens), $0.067 at 1K (1,120), $0.101 at 2K (1,680), and $0.151 at 4K (2,520). The 2.1 table has no 0.5K row. - Search grounding is 5,000 free requests a month, shared across Gemini 3.x models, then $14 per 1,000. One request can trigger more than one billed search. Text or images that search returns are not billed again as input tokens. The paid tier is marked "No" for using the data to improve Google's products.
A 1K picture is about half the image-output price of Nano Banana 2. A 4K picture is not: $0.113 against $0.151, because 2.1 counts 3,780 tokens where Nano Banana 2 counts 2,520. Input tokens cost three times as much, and text or thinking output costs two and a half times as much. The model page says thinking defaults to medium, with minimal and high also available, so a long edit pays the $7.50 text rate on top of the picture.
Two official pages do not agree on the window
The API model page and the 6 October 2026 model card describe different limits. This page does not pick a winner.
- API model page: input limit 131,072 tokens, output limit 32,768. Inputs are text, image, video, and PDF. Output is image and text. Stable id
gemini-nano-banana-2.1. Latest update October 2026. - Model card: based on Gemini 3.6 Flash. Inputs are text and images, with a context window of up to 1 million tokens. It describes image output as a 4K token output and text output as 64K tokens.
- The model card says the Gemini 3.6 Flash knowledge cutoff is March 2026, and that some domains stay at January 2025, in line with the Gemini 3 family.
- Default output size on the model page is 1K, with 2K and 4K also offered. It says tiling artifacts are fixed on the wide ratios 1:4, 4:1, 1:8, and 8:1 at 2K and 4K.
- Up to 14 reference images: character consistency for up to 4 characters, and object fidelity for up to 10 objects.
- Supported on that page: image generation, thinking, search grounding, and the Batch API. Not supported: audio generation, caching, code execution, file search, function calling, Google Maps grounding, the Live API, structured outputs, URL context, Flex inference, and Priority inference.
Where it shows up
The model card lists these channels: the Gemini app, Google AI Studio, the Gemini API, Google Search AI Mode, Google Ads, Google Flow, and Google Stitch. The DeepMind product page that introduces 2.1 still uses the heading Nano Banana 2 for the family. The 2.1 section on that page names three changes: visual design, mask-based editing, and subject consistency.
Scores Google printed, with their own error bars
The model card's October 2026 tables are side-by-side human preference scores, plus one factuality score. They are Google's evals, not a retest here. Each line is Nano Banana 2.1 with thinking, 2.1 without thinking, Nano Banana 2 with thinking, then Nano Banana Pro. A blank was not left blank: every cell below was printed.
- Overall preference: 1050 +/- 14, 1015 +/- 13, 990 +/- 7, 935 +/- 8.
- Infographic design: 1048 +/- 17, 1001 +/- 17, 961 +/- 12, 912 +/- 12. Infographic factuality, a different scale: 0.521, 0.328, 0.179, 0.265.
- General editing: 1026 +/- 12, 980 +/- 15, 938 +/- 11, 939 +/- 10.
- Single-character consistency: 1028 +/- 14, 1021 +/- 14, 981 +/- 10, 991 +/- 9. Multi-character: 1106 +/- 14, 1068 +/- 14, 978 +/- 10, 1011 +/- 10.
- Mask and ink editing: 1049 +/- 15, 1042 +/- 16, 965 +/- 12, 927 +/- 12. Product consistency: 1024 +/- 18, 981 +/- 18, 955 +/- 22, 965 +/- 14.
- Stylization: 1062 +/- 20, 1036 +/- 17, 991 +/- 12, 990 +/- 12. Multi-reference editing: 1066 +/- 22, 1041 +/- 20, 988 +/- 13, 989 +/- 12.
What the model card says still goes wrong
- Small text is often blurry at 1K. Long paragraphs and page-length text still render poorly.
- A character is not always the same person from the input photo to the output. Mask and doodle edits can follow only part of the instruction, and the ink does not always stay.
- Edits sometimes keep the subject's pose from the input. Left and right can get swapped. World knowledge, 3D reasoning, and factuality are still limited. Hallucinations, slowness, and timeouts can still happen.
On safety, the card says child-safety checks met Google's launch thresholds, and that content-safety results were similar to or better than Gemini 3 Flash. For frontier risk it points at Gemini 3.1 Pro and Gemini 3.7 Flash, says those did not reach tracked or critical capability levels, and says 2.1 has no meaningful new capability over them, so Google is confident 2.1 does not either. That is Google's assessment.