File Token Estimator
Drop one or many text-based files to see how many tokens they contain — files are read locally and never uploaded.
Runs in your browser — nothing is uploadedFiles
| File | Characters | Words | Tokens (GPT) |
|---|
How to use File Token Estimator
- Drop files or choose them (text-based formats, up to 50 MB in total).
- See tokens per file and in total for each model family.
- Check whether they fit a model’s context window and what embedding them costs.
Formula & assumptions
- Counts are estimates: the tokenizer approximation was calibrated against OpenAI’s o200k_base (typically within a few percent for English prose and code; up to about 20% off for some other languages). Nothing you paste leaves your browser.
Questions
Which files are supported?
Plain-text formats: .txt, .md, .csv, .json, .jsonl, .html, .xml, .yaml and source code files. PDFs and Word documents need converting to text first (try PDF to Text).
How accurate is it?
Counts are estimates: the tokenizer approximation was calibrated against OpenAI’s o200k_base (typically within a few percent for English prose and code; up to about 20% off for some other languages). Nothing you paste leaves your browser.