AI Reformat Checker

Paste your original data and the version an AI (or any tool) gave back in another format. Every value is matched and checked: rows dropped or added, numbers rounded, units lost, text cut short and columns renamed.

Checker Web & Dev Updated Oct 4, 2026
Learn how this works
How to Use
  1. Paste the original data in the first box: CSV, a spreadsheet paste, JSON, YAML, a Markdown or HTML table, or a “Key: value” list.
  2. Paste the reformatted version you got back in the second box. Both formats are detected for you; set the Format menu if a guess is wrong.
  3. Read the verdict, such as “All 24 values preserved” or “1 changed, 1 row dropped”, and the table of every difference with its before and after.
  4. Tick or untick the options: ignore whitespace, ignore letter case, and treat 1,234 and 1234 as the same value.
  5. Check Show Work for how each side was read, how columns and rows were matched, and why each value was flagged, then press Copy report.
Input
Format: paste the source data
Format: paste what came back
Presets
Differences
Waiting for both versions
Differences appear here, one row per changed value.
Values checked
—
Preserved
—
Changed
—
Rows missing / added
—

Worked Example

Original: a CSV of 5 lab samples with 4 columns (sample, length_mm, ratio, notes), so 5 × 4 = 20 values. Reformatted: the JSON an assistant returned inside a ```json fence, with 4 objects.

Read both sides. The CSV has a comma between fields and the same 4 fields on all 6 lines. The fence is removed, the JSON starts with [ and parses to 4 rows × 4 fields, and A2’s null note is read as blank, matching the empty CSV cell. Numbers keep their written text, so 12.0 stays 12.0 rather than becoming 12.

Match. All 4 column names are identical. The sample column is filled in and unique on both sides, so it is the key: A1, A2, A3 and A5 are found, and A4 is missing.

Compare. The 4 matched rows give 16 cells; 15 are identical. A1’s ratio went from 3.14159 to 3.14, and |3.14159 − 3.14| = 0.00159 ≤ 0.005, half a unit in the second decimal place, so it is flagged as rounded to 2 decimal places.

Verdict: 1 changed, 1 row dropped — 15 of the 20 original values kept; the other 4 were in the dropped row.

A second case: a JSON product list came back as an HTML table with price renamed to Cost and the rows sorted by name. 3 of the 4 prices (29.99, 34.50, 59.00) appear in Cost, a 75% match, so the columns are paired as a rename. The original order A-100, A-101, A-102, A-103 sits at positions 2, 4, 1, 3; the longest run still in order has 2 rows, so 2 rows reordered, and the one real edit, 89.95 → 89.99, is listed on its own.

The common mistake: comparing the two versions row by row. Once A4 is gone, row 4 of the JSON is A5, so a line-by-line diff reports 4 changed cells in row 4 (A4 → A5, 12.25 → 12.0, 1.73205 → 2.23607, blank → final) and a missing row 5, and never says which sample disappeared. Matching rows by a key, or by identical content, reports the one dropped row instead.

Show Work

Paste the original and the reformatted data to see how each was read, matched and compared.

How the Comparison Works

Format detection order
fence → JSON → HTML → Markdown → YAML → key: value → CSV/TSV
JSON starts with { or [; HTML has <tr> and <td>; Markdown has a | --- | line; CSV uses the delimiter (, tab ; |) with the most consistent field count
Column matching
share = |A ∩ B| ÷ max(|A|, |B|) ≥ 60%
Same name first (case, spaces and punctuation ignored); otherwise columns whose distinct values overlap are a rename
Key column
filled + unique on both sides, ≥ 50% of keys found
Names such as id, key, sku, code, email and name are preferred; rows are matched by key
No key: alignment
LCS of identical rows, n × m ≤ 4,000,000
Edited rows between anchors pair up when ≥ 1/3 of their values agree
Rows reordered
moved = matched − LIS(new positions)
LIS is the longest increasing subsequence: the most rows still in their original order
Rounded
|a − b| ≤ ½ × 10−d
d = decimal places of the new value, fewer than the original’s: 3.14159 → 3.14 is 0.00159 ≤ 0.005
Kept, reformatted, changed
kept  |  same meaning  |  changed
Kept: identical, spacing or quotes, 1,234 = 1234 (options). Same meaning: 1.50 → 1.5, 45% → 0.45, a date in a new format, Yes → true. Changed: everything else

Why Reformatting Changes Data

A large language model does not copy your table into a new layout. It writes the answer one token at a time, each token predicted from the text so far, and a token is often a fragment of a word or a few digits of a number. Reproducing 200 rows exactly means getting thousands of predictions right in a row, so occasional slips are expected: a row skipped, a long number with digits swapped, a value “tidied” to fewer decimal places, a unit left off, or a long cell shortened with an ellipsis. Models also have a maximum output length, so a long table can simply stop early or end with a line such as “… and 40 more rows”.

The same kind of damage happens without AI. Spreadsheet programs drop leading zeros from codes such as 00123 and turn text that looks like a date into a date. In 2016 Ziemann, Eren and El-Osta reported in Genome Biology that roughly one in five papers with Excel gene lists in their supplementary files had gene names converted to dates, such as SEPT2 becoming 2-Sep; in 2020 the HUGO Gene Nomenclature Committee renamed genes like SEPT1 and MARCH1 to SEPTIN1 and MARCHF1 partly for this reason. JavaScript’s JSON.parse rounds integers above 253 − 1 = 9,007,199,254,740,991.

That is why the check here compares values, not text: two files can look completely different (a Markdown table and a YAML list) and hold exactly the same data, or look almost the same and differ in one number.

About This Tool

This tool checks that a reformatted copy of some data still holds the same values as the original. It reads both versions whatever their format, turns each into rows and fields, matches columns by name or by shared values, lines up rows by a key column or by identical content, and then compares every cell. Each difference is named: rounded, unit dropped, text truncated, leading zeros lost, case changed, value blanked, row dropped, row added or row moved.

A text diff compares characters, so a CSV against a Markdown table shows every line as changed. This compares values, so the verdict can be “All 24 values preserved” for two files that share no lines. Everything runs in your browser; nothing you paste is uploaded, and pasted HTML is read as text and never rendered.

It suits anyone who asks an AI assistant to convert a table, writes data into documentation, or checks the output of an export script before trusting it.

Related tools: Text Diff Tool, JSON ⇄ CSV Converter, and YAML / JSON Converter.

Frequently Asked Questions

Why would an AI change my data when I only asked it to reformat it?

A language model writes its answer one token at a time, predicting each piece of text rather than copying bytes from your input. Long tables give it many chances to slip: a row can be skipped, 3.14159 can come back as 3.14, “1.2 kg” can lose its unit, and a long description can be shortened with an ellipsis. Very long outputs can also hit the model’s output limit and stop early. Most conversions are fine, which is why the errors are easy to miss.

How are rows matched when some are missing or in a different order?

If a column is filled in and unique on both sides, such as id, sku or name, rows are matched by that key, so a table sorted by name shows “2 rows reordered” instead of every row changed. Without a key, rows are lined up by the longest common subsequence of identical rows, and edited rows between those anchors are paired when at least a third of their values agree. Tables above 4,000,000 row pairs use in-order matching within a 1,000-row window.

Does 1,234 count as the same value as 1234?

By default, yes: “Treat 1,234 and 1234 as equal” is ticked, so “12,480” in a CSV and 12480 in YAML count as kept. Untick it to list them as changes. Values that keep their meaning but not their text, such as 1.50 → 1.5, 45% → 0.45 or 2024-03-05 → March 5, 2024, are counted as changed and marked “only reformatted” in the verdict.

Which formats can it read?

CSV, TSV and spreadsheet pastes (comma, tab, semicolon or pipe), Markdown tables, HTML tables, JSON (arrays of objects, objects of objects, arrays of arrays), simple YAML and “Key: value” lists with bullets, headings or bold labels. A ```json code fence and the chat text around a table are ignored. Only the first table is read, and YAML anchors, tags and multi-line flow collections are not supported.

How much data can it check, and is anything uploaded?

Each side can hold up to 5 MB of text, 50,000 rows, 500 columns and 2.5 million cells. In our tests 50,000 rows × 6 columns compared in about 0.7 seconds. Nothing is uploaded: parsing and comparing run in your browser, and pasted HTML is read as text and never rendered, so scripts in it cannot run.

How do I use the AI Reformat Checker?

Just type your numbers. The answer shows up right away — there is no button to press. Change anything and it updates by itself.

Does it cost anything or need an account?

No. The tool is completely free, there is no account to create, and it keeps working offline after the page first loads.

Is anything I type uploaded?

No. The tool works entirely on your device, so the values you enter never leave your browser.

Common Use Cases

Before pasting an AI table into a report

A 6-row product CSV turned into a Markdown table: the verdict “All 24 values preserved” means you can paste it without reading every cell.

Catching silent rounding

Measurements converted to JSON came back with ratio 3.14 instead of 3.14159: |3.14159 − 3.14| = 0.00159, flagged as rounded to 2 decimal places.

Spreadsheet to documentation

A gear list turned into bullets lost “kg” from 1.2 kg and cut a 38-character description to 18 characters with an ellipsis. Both are listed with before and after.

Checking an export, not just AI

An HTML table from a shop export renamed price to Cost and sorted by name. 3 of 4 prices still match, so the column is paired as a rename and the 89.95 → 89.99 edit stands out.

Generated YAML config

CSV depot data became YAML: “12,480” → 12480 counts as kept, while active → Active is flagged as a letter-case change.

Last updated: