Skip to main content

How Are Generations Counted?

This article explains how pricing works in Pencil, and how various bits of work count against your generations credits. A detailed look at how Pencil counts generations across text, image, video and audio models, and what a typical brief costs end to end.

Written by Michael Whyle

Pricing update - 24 August 2026

We are making three changes to how generations are counted. Together they mean more models to choose from, and lower cost for typical workloads.

  • A lot more models. We have added models from a wider set of providers, including several that perform close to the frontier at a fraction of the cost.

  • More granular rates for every model. Rates now follow what each model costs to run, rather than a flat charge per image or a rounded whole generation. That gives you more control over what a piece of work costs.

  • New agentic pricing for text models. A single per-message rate no longer describes what Scribble does. Text work is now priced across seven size buckets, so a quick edit costs a fraction of a generation and a full campaign brief is priced for what it actually involves.

Overview

In most cases, a generation is every time you interact with an AI model to generate something. Knowing how many generations you have used matters, because most subscriptions include a set number of them.

How you trigger a generation does not change what it costs. You might prompt a model directly, or an agent might generate on your behalf as part of a larger task. Either way, the rules below apply.

Where a cost is a fraction, it is rounded up to one decimal place. Small pieces of work cost a fraction of a generation, not a whole one.

Below are specific details pertaining to:

  • Agentic & Text Work

  • Example Agentic Brief Costs

  • Image Generation

  • Video Generation

  • How the Models Compare

  • Feed Variations

  • AI Voiceovers

  • Scores


Agentic & Text Work

When Pencil’s AI does work for you with a text model — a chat turn with Scribble, a step in a workflow, or an editor action — the amount of work involved can vary a lot. A quick copy tweak might take one small model call. A full campaign brief might take dozens of large ones.

To keep pricing fair, each unit of work is priced by how much work was actually done. After it completes, it is placed into one of seven size buckets, from XS to XXXL, based on two things: how many AI model calls it made, and how much information (tokens) those calls processed. The bucket takes the higher of the two.

The price also depends on the AI model you choose.

Size buckets

Bucket

AI model calls

Total tokens processed

XS

≤1

<50K

S

2–3

50–300K

M

4–10

300K–1M

L

11–30

1–3M

XL

31–100

3–10M

XXL

101–300

10–30M

XXXL

>300

>30M

Note: the bucket is whichever tier is higher. Work that makes 5 calls (M) but processes 2M tokens (L) lands in L.


Generations per unit of work, by model and bucket

Model name

XS

S

M

L

XL

XXL

XXXL

Anthropic: 9 models

0.4

6.9

14.2

34.5

63.3

138.0

299.0

Fable 5

0.7

13.3

27.2

66.4

121.8

265.6

575.4

Claude Opus 5

0.6

10.6

21.6

52.7

96.5

210.5

456.0

Claude Opus 4.8

0.5

9.3

19.0

46.3

84.9

185.2

401.2

Claude Opus 4.6

0.5

9.3

19.0

46.3

84.9

185.2

401.2

Claude Opus 4.7

0.4

7.7

15.7

38.2

70.1

152.8

331.0

Claude Sonnet 4.6

0.3

4.1

8.3

20.2

37.0

80.7

174.7

Claude Sonnet 4.5

0.3

4.1

8.3

20.2

37.0

80.7

174.7

Claude Sonnet 5

0.2

2.7

5.5

13.5

24.7

53.8

116.5

Claude Haiku 4.5

0.1

1.4

2.9

7.0

12.8

27.8

60.3

OpenAI: 10 models

0.1

2.0

4.1

10.0

18.2

39.7

85.9

GPT-5.4

0.2

3.9

7.9

19.1

35.1

76.4

165.6

GPT-5.6 Sol

0.2

3.3

6.6

16.2

29.6

64.6

139.8

GPT-5.5

0.2

3.3

6.6

16.2

29.6

64.6

139.8

GPT-5.2

0.2

2.7

5.4

13.2

24.1

52.6

113.9

GPT-5.3 codex

0.2

2.7

5.4

13.2

24.1

52.6

113.9

GPT-5.6 Terra

0.1

1.7

3.3

8.1

14.8

32.3

69.9

GPT-5.4 mini

0.1

1.1

2.3

5.5

10.1

21.9

47.5

GPT-5.1

0.1

1.1

2.2

5.4

9.9

21.5

46.5

GPT-5.4 nano

0.1

0.4

0.7

1.7

3.1

6.8

14.6

GPT-5.6 Luna

0.1

0.2

0.4

0.9

1.5

3.3

7.0

Zhipu: 1 model

0.1

1.2

2.3

5.6

10.3

22.4

48.4

GLM-5.2

0.1

1.2

2.3

5.6

10.3

22.4

48.4

Google: 9 models

0.1

1.2

2.3

5.6

10.2

22.2

48.1

Gemini 3.1 Pro Preview

0.2

2.7

5.5

13.3

24.3

52.9

114.5

Gemini 3.5 Flash

0.2

2.2

4.4

10.7

19.5

42.5

92.0

Gemini 2.5 Pro TTS

0.1

1.9

3.9

9.5

17.4

38.0

82.2

Gemini 2.5 Pro

0.1

1.2

2.5

6.0

10.9

23.7

51.2

Gemini 2.5 Flash

0.1

0.9

1.9

4.5

8.3

17.9

38.8

Gemini 3.6 Flash

0.1

0.6

1.3

3.0

5.5

12.0

25.9

Gemini 3 Flash

0.1

0.5

0.9

2.1

3.7

8.1

17.5

Gemini 2.5 Flash Lite

0.1

0.2

0.4

1.0

1.7

3.7

8.0

Gemini 3.1 Flash Lite

0.1

0.1

0.2

0.4

0.6

1.3

2.7

DeepSeek: 1 model

0.1

1.1

2.1

5.1

9.2

20.1

43.5

DeepSeek V4 Pro

0.1

1.1

2.1

5.1

9.2

20.1

43.5

Moonshot: 1 model

0.1

0.8

1.6

3.9

7.1

15.3

33.2

Kimi K2.6

0.1

0.8

1.6

3.9

7.1

15.3

33.2

Alibaba: 1 model

0.1

0.4

0.7

1.7

3.0

6.5

14.0

Qwen 3.6 Plus

0.1

0.4

0.7

1.7

3.0

6.5

14.0

MiniMax: 3 models

0.1

0.2

0.4

0.9

1.7

3.6

7.8

MiniMax M2.7-highspeed

0.1

0.3

0.6

1.4

2.5

5.3

11.5

MiniMax M3

0.1

0.2

0.3

0.7

1.3

2.8

5.9

MiniMax M2.7

0.1

0.2

0.3

0.7

1.3

2.8

5.9

Average across all 35 models

0.2

2.8

5.6

13.7

25.1

54.7

118.4

Note: provider rows show the average across that provider’s models. After any turn or workflow step completes, open its cost menu (alongside content provenance) to see the bucket, the generations charged, and the models used. You will not be charged for failed generations.


Example Agentic Brief Costs

These examples show what typical briefs cost end to end, using GPT-5.5 for the agentic work plus media generations at the standard rates below. Your actual cost depends on the models chosen and the size of the work.

Brief type

Typical bucket

Agentic work (gens)

Media gens

Total (gens)

Per asset (gens)

Copy & messaging refresh
10 caption or headline variants from one brief

S

3.3

3.3

0.4

Format adaptation / resize
8 formats resized and rebalanced from one master

M

6.6

6.6

0.9

New static concepts
4 new static concepts, ~2 images generated each

L

16.2

19.2

35.4

4.5

Video cutdowns
3 six-second cutdowns from a master video

L

16.2

16.2

5.4

Short film
20-second film — 4 shots, 2 six-second takes per shot, 48 seconds generated

L

16.2

48.0

64.2

64.2

Full campaign pack
12 mixed statics and cutdowns from one brief

XL

29.6

19.2

48.8

4.1

Average across these six briefs

14.7

14.4

29.1

13.2

Note: media assumes Nano Banana Pro for stills and Google Veo 3.1 for video — the highest-rated model in each case, and also the most expensive. Your model choice moves these totals a long way: the same 6-second film generated on Kling AI VIDEO 2.5 Turbo costs 8.4 media generations rather than 48.0, taking the brief from 64.2 to 24.6. Adapting, resizing or cutting down an existing master does not generate new media, so those briefs carry no media generations — only the agentic work is charged.


Image Generation

Image generation is priced per image, at a rate set by the model you choose. Every image generated counts, whatever its resolution or quality.

Model Name

Generations

Adobe: 2 models

3.2

/ image

Adobe Firefly 5

3.2

/ image

Adobe Firefly V4 Standard

3.2

/ image

Getty: 1 model

2.1

/ image

Getty Images (Generative)

2.1

/ image

OpenAI: 2 models

1.1

/ image

OpenAI GPT Image 1.5

1.1

/ image

OpenAI GPT Image 2

1.0

/ image

Google: 8 models

0.7

/ image

Nano Banana Pro (Gemini 3 Pro Image)

2.4

/ image

Nano Banana 2 (Gemini 3.1 Flash Image)

0.7

/ image

Google Imagen 4 Ultra

0.6

/ image

Imagen 4 Standard

0.4

/ image

Nano Banana (Gemini 2.5 Flash Image)

0.4

/ image

Nano Banana 2 Lite

0.4

/ image

Google Imagen 3

0.3

/ image

Imagen 4 Fast

0.2

/ image

Stability: 2 models

0.5

/ image

Stable Diffusion 3 / 3.5 Large

0.7

/ image

Stable Diffusion (SDXL / Core)

0.3

/ image

ByteDance: 3 models

0.5

/ image

Seedream V5 Pro

0.7

/ image

Seedream 4.5

0.4

/ image

Seedream V5 Lite

0.4

/ image

xAI: 2 models

0.4

/ image

Grok Imagine Image v2.0

0.6

/ image

Grok Imagine Image

0.2

/ image

Bria: 1 model

0.4

/ image

Bria 3.2

0.4

/ image

Alibaba: 2 models

0.4

/ image

Qwen Image 2

0.4

/ image

Qwen Image 3

0.3

/ image

Black Forest Labs: 1 model

0.3

/ image

Flux 2 Pro

0.3

/ image

Kuaishou: 4 models

0.3

/ image

Kling AI IMAGE 2

0.3

/ image

Kling AI IMAGE 2.1

0.3

/ image

Kling AI IMAGE 3.0

0.3

/ image

Kling AI IMAGE 3.0 Omni

0.3

/ image

MiniMax: 1 model

0.1

/ image

MiniMax image-01

0.1

/ image

Average across all 29 models

0.8

/ image

Note: post generation actions will have separate cost, e.g. 1 edit = 1 generation on that model’s rate.



​Video Generation

Video generation is priced per second of generated video, at a rate set by the model you choose, but with an important rule:

Where the cost is a fraction, the generations are rounded up to one decimal place.

Model Name

Generations

Adobe: 1 model

5.0

/ sec

Adobe Firefly Video (Video1 Standard)

5.0

/ sec

ByteDance: 4 models

3.7

/ sec

Seedance 2.0

6.5

/ sec

Seedance 2.5

4.5

/ sec

Seedance 2.0 Fast

2.3

/ sec

Seedance 2.0 Mini

1.5

/ sec

OpenAI: 2 models

3.0

/ sec

Sora 2 Pro

5.0

/ sec

Sora 2

1.0

/ sec

Black Forest Labs: 1 model

2.8

/ sec

FLUX 3

2.8

/ sec

Alibaba: 1 model

2.7

/ sec

Happy Horse

2.7

/ sec

Stability: 1 model

2.0

/ sec

Stable Video (image-to-video)

2.0

/ sec

Google: 4 models

1.9

/ sec

Google Veo 3.1

4.0

/ sec

Google Veo 3.1 Fast

1.5

/ sec

Gemini Omni Video (Omni Flash)

1.0

/ sec

Google Veo 3.1 Lite

0.8

/ sec

Kuaishou: 4 models

1.6

/ sec

Kling AI VIDEO 2.1

2.8

/ sec

Kling AI VIDEO 3.0

1.4

/ sec

Kling AI VIDEO 3.0 Omni

1.4

/ sec

Kling AI VIDEO 2.5 Turbo

0.7

/ sec

xAI: 2 models

1.6

/ sec

Grok Imagine Video v1.5

2.4

/ sec

Grok Imagine Video

0.7

/ sec

Runway: 4 models

1.3

/ sec

Runway Aleph 2.0

2.8

/ sec

Runway Gen-4.5

1.2

/ sec

Runway Gen-3 Alpha Turbo

0.5

/ sec

Runway Gen-4 Turbo

0.5

/ sec

Lightricks: 2 models

1.3

/ sec

LTX 2.5

1.7

/ sec

LTX 2.3

0.8

/ sec

MiniMax: 1 model

1.0

/ sec

MiniMax-H3

1.0

/ sec

Average across all 27 models

2.2

/ sec

Examples:

  • A 5-second Runway Gen-4 Turbo output = 5 × 0.5 = 2.5 generations

  • A 3-second Google Veo 3.1 output = 3 × 4.0 = 12.0 generations

  • A 10-second Gemini Omni Video output = 10 × 1.0 = 10.0 generations

Note: all video rates above are for 1080p output, except Gemini Omni Video, which generates at 720p and has no 1080p mode. Where a model publishes several resolutions we show the 1080p tier, or the nearest tier the vendor publishes, so the rates are comparable model to model. Generating at 720p costs less on most models and at 4K costs more; the cost menu on any completed job shows the rate that was actually applied.

Note: you will not be charged for failed generations. Kling models must be enabled by a Workspace Admin before use. Gemini Omni Flash is in preview from Google, so quality and behaviour may change.


Priority Capacity

Models listed in the platform as [Model name] [Priority Capacity] run on throughput we reserve with the provider in advance. They return at lower latency and stay available during peak demand periods, when shared capacity can queue. There is nothing to reserve or arrange: pick the Priority Capacity version from the model list whenever you want it, and the standard version whenever you do not. It costs one extra generation per unit.

Model

Standard

Priority Capacity

Unit

Charged per image

Adobe Firefly 5

3.2

4.2

/ image

Nano Banana Pro (Gemini 3 Pro Image)

2.4

3.4

/ image

Nano Banana 2 (Gemini 3.1 Flash Image)

0.7

1.7

/ image

Nano Banana 2 Lite

0.4

1.4

/ image

Bria 3.2

0.4

1.4

/ image

Charged per second

Google Veo 3.1

4.0

5.0

/ sec

Topaz AI Video Enhance

1.6

2.6

/ sec

Google Veo 3.1 Fast

1.5

2.5

/ sec

Runway Gen-4.5

1.2

2.2

/ sec

Note: Priority Capacity can be chosen per job and dropped again on the next one. There is no commitment, no lead time and no minimum. Every other model runs on shared capacity, which is the default and is priced as published. Available on the models listed here, which cover Google video and image, Adobe Firefly, Runway, Topaz and Bria.


How the Models Compare

Rates differ a great deal between models, and so does output quality. These charts put the two together: how many generations a model charges against its standing on the public Arena leaderboards, which are decided by blind head-to-head votes from the community. Cheaper is to the left, better is higher, so the best value sits toward the top left.

How to read the line graphs: the dashed line joins the best-quality model available at each price, so anything sitting below it is beaten on both counts by something cheaper. A triangle means fewer than 5,000 votes have been cast on that model, so its position is less settled.

Note: the three Arena boards are scored independently, so an Elo of 1450 in video and 1450 in text are not comparable to each other. Each chart must be read on its own. Quality scores are from the public Arena leaderboards, read on 24 August 2026, and are not produced by Pencil.


AI Voiceovers, Sound & Music

Speech, dubbing, sound effects, transcription and music are charged per model, the same way images and video are. Rates below are grouped by the unit the work is measured in.

Model

Generations

Unit

Charged per minute of audio

ElevenLabs 5 models

1.0

/ min of audio

ElevenLabs AI Dubbing

2.1

/ min of audio

ElevenLabs Music

1.5

/ min of audio

ElevenLabs Voice Isolation

0.8

/ min of audio

ElevenLabs SFX

0.5

/ min of audio

ElevenLabs Scribe v2 (STT)

0.1

/ min of audio

Google 2 models

0.3

/ min of audio

Gemini TTS 2.5 (2.5 Pro TTS)

0.3

/ min of audio

Gemini TTS 3.1

0.3

/ min of audio

Charged per 1,000 characters of script

ElevenLabs 3 models

0.5

/ 1,000 characters

ElevenLabs Eleven v3

0.6

/ 1,000 characters

ElevenLabs Multilingual v2

0.6

/ 1,000 characters

ElevenLabs Flash v2.5

0.3

/ 1,000 characters

Google 1 model

0.3

/ 1,000 characters

Google Chirp 3 HD

0.3

/ 1,000 characters

Charged per generated clip

Google 2 models

0.5

/ 30-sec clip

Lyria 2

0.6

/ 30-sec clip

Lyria 3

0.4

/ 30-sec clip

Average across all 13 models

0.7

Note: per-minute rates are charged on the duration of the rendered audio, rounded up to the nearest tenth of a generation. Speech synthesis charged per clip covers one rendered voiceover. Transcription is charged on the length of the source audio, not the transcript. Music models are being brought online through the autumn; the rate shown is the one they will launch on.


Upscaling, Restoration & Avatars

These models operate on an asset that already exists rather than generating a new one. Upscale and restore are priced per megapixel for stills and per second for footage; avatar and lipsync are priced per second of output.

Model Name

Generations

Bria 1 model

1.4

Bria Videos (VRMBG 2.0)

1.4

/ sec

Topaz Labs 8 models

0.9

Topaz AI Video Enhance

1.6

/ sec

Topaz Proteus Natural

1.6

/ sec

Topaz Starlight Precise 2.5

1.6

/ sec

Topaz High Fidelity V2

0.4

/ 24 MP

Topaz Recovery V2

0.4

/ 4 MP

Topaz Redefine

0.4

/ 4 MP

Topaz Standard MAX

0.4

/ 4 MP

Topaz Wonder

0.4

/ 4 MP

HeyGen 3 models

0.6

HeyGen Avatar 5 Digital Twin

0.7

/ sec

HeyGen V3 Lipsync Precision

0.7

/ sec

HeyGen V3 Lipsync Speed

0.4

/ sec

MiniMax 1 model

0.4

MiniMax-H3-Regeneration (768P to 2K)

0.4

/ sec

Average across all 13 models

0.8

Note: a megapixel rate is charged on the output size, not the input, so upscaling a small image to a large one costs more than restoring one at its original size.



Feed Variations

Feed work is charged on what is generated, not on how many formats or cells a job touches. Versioning one creative across many formats stays a single generation, which makes it the cheapest production work on the platform.

Action

Generations

How it is counted

Feed variation

1.0

1 gen per variation generated

One creative, many formats

1.0

1 gen regardless of format count

Feed Sync

0.1

1 gen per ten variations

AI tool used inside a feed

1.0

1 gen per cell actioned

After Effects import

0.1

1 gen per 10 sec of rendered output

Note: Feed Sync is charged at a tenth of the rate of individually generated variations, so routing bulk variation work through Feed Sync is materially cheaper. After Effects imports are priced on the rendered duration of the output, not the length of the project file.


Scores

Default Pencil scores are included and consume nothing. Custom Scores set to Automatic run whenever a work item meeting their criteria is saved, and each run is charged. Scoring agents that call an external data provider cost more per request.

Score

Generations

How it is counted

Custom Score set to Automatic

1.0

1 gen per run

Default Pencil scores

0

No generations consumed

GWI Insight Agent

4.0

4 gens per request

Share of Model score

2.0

2 gens per request

Note: scores are charged on top of the work item itself. A generated image meeting the criteria of three Custom Scores set to Automatic consumes four generations in total — one for the image, one for each score. Setting a Custom Score to run manually rather than automatically puts you in control of when it is charged.


Tracking Generations

You can view your remaining generations in the counter at the bottom right of your workspace, per-user usage in any workspace’s Manage Team menu, and per-turn cost in the cost menu on any Scribble turn or workflow step.


Did this answer your question?