File Token Estimator

Drop one or many text-based files to see how many tokens they contain — files are read locally and never uploaded.

Runs in your browser — nothing is uploaded

How to use File Token Estimator

  1. Drop files or choose them (text-based formats, up to 50 MB in total).
  2. See tokens per file and in total for each model family.
  3. Check whether they fit a model’s context window and what embedding them costs.

Formula & assumptions

  • Counts are estimates: the tokenizer approximation was calibrated against OpenAI’s o200k_base (typically within a few percent for English prose and code; up to about 20% off for some other languages). Nothing you paste leaves your browser.

Questions

Which files are supported?

Plain-text formats: .txt, .md, .csv, .json, .jsonl, .html, .xml, .yaml and source code files. PDFs and Word documents need converting to text first (try PDF to Text).

How accurate is it?

Counts are estimates: the tokenizer approximation was calibrated against OpenAI’s o200k_base (typically within a few percent for English prose and code; up to about 20% off for some other languages). Nothing you paste leaves your browser.