Token Counter

How many tokens is my prompt?

Paste any text. You see every token, the count for GPT-5, GPT-4o, GPT-4 and open models, what the prompt costs, and which wordy phrases to cut.

Runs 100% in your browser. Your text is never uploaded.

Nothing to count yet.

Paste a prompt above, or load the sample: a 5-line support-bot prompt that is 70 tokens for GPT-4o and 61 after the suggested cuts.

How tokenization works

A language model never sees letters. Before your prompt reaches it, a tokenizer cuts the text into pieces from a fixed vocabulary and swaps each piece for a number, its token ID. Billing, context limits and speed are all counted in these pieces, which is why a token counter matters more than a word count.

From text to token IDs The sentence "Tokenizers split unbelievable words." is split by a pattern into 5 chunks, byte pair merges turn them into 6 tokens (Token, izers, split, unbelievable, words, period), and each token becomes an o200k_base ID: 4421, 24223, 12648, 83614, 6391, 13. 1. Your text2. Split by a pattern (words, spaces, digits, punctuation)3. Merge bytes into known pieces (byte pair encoding)4. Look up each piece's ID Tokenizers split unbelievable words. Tokenizers ␣split ␣unbelievable ␣words . Token izers ␣split ␣unbelievable ␣words . 4421242231264883614639113 5 chunks6 tokenso200k_base
Real output of the o200k_base tokenizer (GPT-4o, GPT-4.1, GPT-5). The ␣ marks a space that belongs to the token.

Spaces are part of the token

Most tokenizers glue the space to the word after it. In o200k_base, " unbelievable" with its leading space is 1 token, while "unbelievable" at the start of a line is 3 (un, bel, ievable). That is why the same word can cost a different amount in different places.

Every model has its own vocabulary

o200k_base has about 200,000 pieces and cl100k_base about 100,000; Llama 3, Qwen3, DeepSeek and Mistral each ship their own. A bigger vocabulary usually needs fewer tokens for the same text, especially outside English. Try the unicode in your own prompt and switch tokenizers above.

Why cutting words saves money

You pay per input token on every request, so a system prompt sent a million times multiplies every wasted phrase. "In order to" is 3 tokens where "to" is 1. The suggestions above swap five such phrases, trim extra whitespace, and then count again to show the real saving.

Questions

How many tokens is my prompt?

Paste it into the box at the top of this page. The count updates as you type, for every tokenizer listed: o200k_base (GPT-5.x, GPT-4.1, GPT-4o, o3, o4-mini), cl100k_base (GPT-4, GPT-3.5 Turbo) and, once you pick them, Llama 3, Qwen3, DeepSeek-V3 and Mistral. OpenAI's own rule of thumb is about 4 characters of English per token, but this page counts exactly with the real tokenizer instead of estimating.

Which tokenizer do GPT-5, GPT-4o and GPT-4 use?

According to the model table in OpenAI's tiktoken library (version 0.14.0), GPT-5 and its point releases, GPT-4.1, GPT-4o, o1, o3 and o4-mini use o200k_base. GPT-4, GPT-4 Turbo, GPT-3.5 Turbo and the text-embedding-3 models use cl100k_base. GPT-6 models are not in that table yet, so this page does not claim a count for them.

Can it count Claude or Gemini tokens?

No. Anthropic and Google do not publish the tokenizers for their current models, so any count made in a browser would be a guess. Both offer a token counting endpoint in their APIs if you need an exact number.

Is my text uploaded anywhere?

No. The tokenizers run in your browser. The page downloads the tokenizer files from this site and never sends your text back. A share link stores the text after the # in the address, and browsers do not send that part of a URL to the server.

Why does the same text give different counts for different models?

Each model family has its own vocabulary of text pieces. A larger vocabulary, like o200k_base with about 200,000 entries, usually covers common words and non-English scripts in fewer pieces than cl100k_base with about 100,000. For plain English the counts are often close; languages other than English and emoji are where they differ most.

Does the count include the chat formatting the API adds?

No. It counts the text you paste and nothing else. Chat APIs wrap every message in a few extra formatting tokens, and some open models (Llama 3 and Mistral, for example) add a start-of-text token, so a real request is a little larger than this count.

How is the cost calculated?

Tokens multiplied by each model's standard input price per million tokens, as listed on OpenAI's API pricing page on the date shown under the table. Output tokens, cached input discounts, batch pricing and regional uplifts are not included. Open models have no single price because each host sets its own, so only their counts are shown.

How do I reduce the tokens in a prompt?

Use the suggestions panel. It applies the same rules as the Token-Visualizer command-line tool: it swaps five wordy phrases for short ones (in order to, due to the fact that, at this point in time, for the purpose of, in the event that), flags words repeated more than three times, lines over 40 tokens and a low characters-per-token ratio, and removes extra spaces and blank lines. It then tokenizes the shortened text again and shows the exact saving.

Is there a command-line version?

Yes. Token-Visualizer is a free, open-source Python tool on GitHub that prints the same token split, line ranking and suggestions in a terminal, and can load any Hugging Face tokenizer by model ID.