Mojibake Repair Lab

Paste garbled text like “café”, “’”, “Привет”, “縺薙s縺ォ縺。縺ッ” or “ÄãºÃ” and get the original back: café, ’, Привет, こんにちは, 你好. The lab works out which character sets were mixed up (UTF-8, Windows-1250/1251/1252, Mac Roman, DOS, Shift_JIS, GBK, Big5, EUC-KR and more), undoes double-encoding layer by layer, fixes UTF-16 files opened as 8-bit text, decodes stray HTML entities and %-escapes, and shows how any text would break. It runs entirely in your browser.

Tool Text & Encoding Updated Oct 3, 2026
How to Use
  1. Paste the garbled text into the box, or pick an example. The repair runs as you type; the box opens with “café — it’s a naïve résumé”.
  2. Read the diagnosis and the readouts: which character set the text was saved in, which one it was wrongly opened as, how many layers deep the mistake goes and how many sequences were repaired.
  3. If the lab offers several possible readings (common with old Asian and Cyrillic text), pick the one that makes sense, or open Set the character sets by hand and choose both yourself.
  4. Check the working in Show Work below the tool: each broken sequence turned back into bytes and decoded correctly, layer by layer.
  5. Copy the repaired text. Use the Simulate tab to see how clean text breaks under 14 common mix-ups, which helps you recognise what you are looking at.
Input
Runs locally; nothing is uploaded.
Set the character sets by hand
Repair
Diagnosis
—
Layers
—
Repairs
—
Lost bytes
—

Worked Example

café → café. The letter é is U+00E9, which UTF-8 stores as the two bytes C3 A9. A program that assumes Windows-1252 maps each byte to its own character: C3 is à and A9 is ©, so “café” shows as “café”. The repair runs the same steps backwards: turn é back into the Windows-1252 bytes C3 A9, then decode those bytes as UTF-8 to get é. The box opens with “café — it’s a naïve résumé”: 34 characters in which the lab finds 6 broken sequences, one layer deep, and returns the 26-character original. The dash — is E2 80 94, which Windows-1252 shows as â, € and ”.

Привет → Привет. The six Cyrillic letters of Привет are 12 bytes in UTF-8 (D0 9F D1 80 D0 B8 D0 B2 D0 B5 D1 82). Read as Windows-1251, which is itself a Cyrillic character set, every byte becomes a Cyrillic letter or symbol, so the result looks Russian but is nonsense: D0 9F (П) shows as Р and Џ. The opposite mistake, Windows-1251 bytes CF F0 E8 E2 E5 F2 opened as Western, gives “Ïðèâåò”.

The common mistake: converting the file to UTF-8 again. If café is already on screen, “converting to UTF-8” saves the mojibake itself: é becomes the four bytes C3 83 C2 A9, which show as “café”. Each attempt adds a layer, and é grows from 1 character to 2 to 4. The fix is the reverse: decode as the wrong character set, re-read as UTF-8, once per layer. This lab undoes both layers and reports 2.

Show Work

Paste garbled text above to see each broken sequence turned back into bytes and decoded.

Formulas and Reference

How mojibake is made
shown = decodewrong(encoderight(text))
How it is repaired
text = decoderight(encodewrong(shown))
é → é
UTF-8 opened as Windows-1252 (C3 A9)
РџСЂ → Пр
UTF-8 opened as Windows-1251 (D0 9F D1 80)
Ïðèâåò → Привет
Windows-1251 opened as Western
縺薙s → こん
UTF-8 opened as Shift_JIS (E3 81 93 E3 82 93)
é → é
Double-encoded: two layers
锟斤拷
�� (EF BF BD EF BF BD) read as GBK: unrecoverable

Finding Out Which Character Sets Were Mixed Up

The hard part is knowing which two character sets were involved. The lab tries every likely combination (UTF-8 and the Windows code pages for Western, Central European, Cyrillic, Greek, Turkish, Hebrew, Arabic, Baltic, Vietnamese and Thai text, Mac Roman, the DOS and Windows console code pages, KOI8-R, Shift_JIS, EUC-JP, GBK, Big5 and EUC-KR) and scores each result for how much it looks like real writing. UTF-8 repairs are self-checking, because random bytes almost never form valid UTF-8, so they are applied automatically. The older 8-bit character sets reuse the same byte values for different alphabets, so when two readings are equally plausible you see both and choose. If the same mistake happened more than once, the layers are peeled off one at a time.

A UTF-16 file opened as ordinary text shows every other byte as a NUL, or turns English into CJK-looking characters when the byte order is wrong; both are detected and decoded. Some bytes cannot survive a wrong decode at all: Windows-1252 has no character for 81, 8D, 8F, 90 and 9D, so a closing curly quote can shrink to †(restored as a marked guess), and anything already shown as �, or as 锟斤拷, is gone for good. Text that is escaped rather than garbled, such as é, %C3%A9, \u00e9 or quoted-printable =C3=A9, is spotted too and decoded on request.

Where Mojibake Comes From

The word is Japanese: mojibake (文字化け) joins moji, “character”, and bake, “to change form”. Japan had good reason to name it, because Japanese computers used three incompatible encodings at once: Shift_JIS on PCs, EUC-JP on Unix and ISO-2022-JP in email. Russia had the same problem with KOI8-R, Windows-1251 and the DOS code page 866, so the same Cyrillic letters could arrive in three different byte forms.

Every one of these character sets used the same 128 ASCII codes for English and gave the upper half of the byte (128 to 255) to its own alphabet. Nothing in a plain text file says which one was meant, so the reader had to guess. UTF-8, designed by Ken Thompson and Rob Pike in September 1992, ended the guessing for new text by giving every character in Unicode one byte sequence, and it now encodes the large majority of web pages. The mojibake that remains comes from the boundary: old files, databases and spreadsheets that still assume Windows-1252 or another code page and read UTF-8 bytes one at a time.

About This Tool

Mojibake Repair Lab diagnoses garbled text, repairs it through up to four layers of mis-decoding, and shows every repaired sequence with its bytes and code points. It covers 20 character sets, from UTF-8 and the Windows code pages to Shift_JIS, GBK, Big5 and EUC-KR, plus UTF-16 files opened as 8-bit text and escaped text. Where a repair is a guess, it says so; where two readings are equally likely, it shows both rather than picking one silently. The Simulate tab runs the mistakes forwards, so you can match the garbage in front of you to the mix-up that made it.

Everything runs in your browser using its built-in decoders; nothing is uploaded, logged or sent to a server or AI. It is for anyone cleaning up a CSV export, an old database, a mailbox or a log file.

Related tools: Unicode Lab, Hex Editor & Viewer, and File Inspector.

Frequently Asked Questions

What is mojibake?

Mojibake is garbled text that appears when bytes are decoded with the wrong character set. A character like “é” is stored in UTF-8 as two bytes, C3 A9; a program that reads them as Windows-1252 shows the two characters “é”. The original bytes are still there, just labelled wrong, so the text can usually be recovered by turning the characters back into bytes and decoding them correctly. That is what this tool does.

Which mix-ups can it fix?

UTF-8 opened as Windows-1252/Latin-1, Windows-1250 (Central European), Windows-1251 (Cyrillic), Greek, Turkish, Hebrew, Arabic, Baltic, Thai, Mac Roman, the DOS/Windows console (CP437/CP866), KOI8-R, Shift_JIS, EUC-JP, GBK, Big5 and EUC-KR, and the reverse, where old Cyrillic, Greek, Hebrew, Arabic, Thai, Japanese, Chinese or Korean text was opened as Western. It also undoes double and triple encoding (up to 4 layers), repairs UTF-16 files opened as 8-bit text (spaced-out letters or NUL bytes), and decodes stray HTML entities, %-escapes, \u escapes and quoted-printable bytes.

How does it know which character sets were mixed up?

It tries every likely pair, decodes the text each way, and scores each result for how much it looks like real writing: no impossible byte patterns, no letters from two alphabets glued into one word, common rather than rare Chinese, Japanese and Korean characters, correct Thai spelling, and so on. UTF-8 is self-checking, because random bytes almost never form valid UTF-8, so those repairs are applied automatically. Old 8-bit character sets reuse the same bytes for different alphabets, so when two readings are equally plausible the lab shows both and lets you choose. Nothing is uploaded: it all runs on your browser’s own decoders.

What does “double-encoded” mean?

The same mistake happened twice: text that was already mojibake was saved as UTF-8 and mis-read again. “é” (C3 A9) shows as “é”; saving that as UTF-8 gives C3 83 C2 A9, which shows as “é”. One character became 2, then 4. The lab peels the layers off one at a time (up to four) and shows each one in Show Work.

Why can’t some text be fully repaired?

If a character was already replaced with “�” (the replacement character), or a byte was dropped, that information is gone. Windows-1252 has no character for the bytes 81, 8D, 8F, 90 and 9D, so some programs drop them: a closing quote ” (E2 80 9D) then shows as just “—, and the lab restores it as a marked best guess. Reading UTF-8 as Big5 or EUC-KR destroys many bytes, so those repairs are partial. “锟斤拷” is the classic unrecoverable case: it is two replacement characters (EF BF BD EF BF BD) read as GBK.

How do I use the Mojibake Repair Lab?

Just type your numbers. The answer shows up right away — there is no button to press. Change anything and it updates by itself.

Does it cost anything or need an account?

No. The tool is completely free, there is no account to create, and it keeps working offline after the page first loads.

Is anything I type uploaded?

No. The tool works entirely on your device, so the values you enter never leave your browser.

Common Use Cases

CSV opened in Excel

Fix names that came out as “José García”, or Cyrillic that came out as “Привет” (12 bytes of Привет read one byte at a time).

Old database or website

Recover text stored in Windows-1251, Shift_JIS or GBK that a modern page displays as accented garbage, such as Ïðèâåò for Привет.

Double-encoded exports

Undo “é” in two layers: back to “é”, then to “é”.

Email, RSS and API content

Undo smart quotes that became “ and ’, and decode leftover =C3=A9 or é escapes.

Log files and the console

Turn “caf├⌐” from a Windows terminal (CP437), or a UTF-16 log full of NUL bytes, back into readable text.

Last updated: