PDFIntact

PDF에서 표와 본문을 그대로 꺼냅니다

Will this PDF
come out clean?

The symptom

This is what copying gives you

Common in papers and government documents. The page looks perfectly fine on screen, so there is no way to spot it by eye.

As shown on screen
A Japanese government circular — perfectly readable on screenThis example is a Japanese document, where this fault is most common. The same cause and the same fix apply to documents in any language.
Copy the title and you get

ᆅ᪉බඹᅋయ࡟࠾ࡅࡿ᭩㠃つไࠊᢲ༳ࠊᑐ㠃つไࡢぢ┤ࡋ࡟ࡘ࠸࡚

The actual text

地方公共団体における書面規制、押印、対面規制の見直しについて

□ は、その文字を表示できる書体が手元に無いという印です。 化けた先が20以上の文字体系にばらけるため、こうなります。この通知PDFは上のサンプルからそのまま試せます。

A real result

This is what conversion gives you

The output of a real PDF put through the service. Not a hand-made example.つまみを動かすと、原本と変換結果を重ねて見比べられます。

原本 PDF総務省統計局「日本統計年鑑」13-5 鉄道輸送量のページ画像(変換前の原本)
変換結果
年度貨物輸送量
貨物数量 (1,000トン)貨物トンキロ (100万トンキロ) 1)
コンテナ車扱コンテナ車扱
平成 30 年42,32123,05019,27119,36917,7241,645
令和 元 年42,66023,50619,15419,99318,3821,610
239,12421,27317,85018,34016,8381,502
年度旅客輸送量
旅客数量 (100万人)旅客人キロ (100万人キロ) 2)
定期定期外定期定期外
平成 30 年25,26914,62710,642441,614212,055229,559
令和 元 年25,19014,79710,392435,063213,511221,552
217,67011,2526,418263,211160,549102,662
JR
平成 30 年9,5565,8173,739277,670113,177164,493
令和 元 年9,5035,8763,627271,936113,907158,029
26,7074,6082,099152,08487,86864,216
# 新幹線1564211434,9363,64231,294
民鉄
(JR以外)
平成 30 年15,7148,8106,903163,94498,87865,066
令和 元 年15,6878,9216,765163,12699,60463,523
210,9636,6444,319111,12772,68038,447
年度索道旅客輸送量
旅客数量 (1,000人)旅客収入 (100万円)
普通索道
3)
特殊索道
4)
普通索道
3)
特殊索道
4)
平成 30 年295,19450,235244,95970,65827,51443,145
令和 元 年239,59445,441194,15363,01725,96437,053
2200,06923,843176,22643,34313,81229,531

このページの実測値: 数値125個中124個が一致(99.2%)。1個だけずれています。

How we handle accuracy

We never mix what we read with what we guessed

Every value we output carries one of three labels. Text and tables are read from the document and cross-checked against its own text layer where one exists. Charts without printed numbers are estimated from the shape of the bars and lines, and are always marked as estimates. Anything we could not read is left out rather than guessed — because a wrong number that looks confident is worse than a gap.

グラフは弱点です。数値のラベルが図の中に書かれていないグラフは、 曲線や棒の高さからの推定になります。そのため確定値としては扱わず、推定値と分かる形で表示します。 グラフの数値を確定値として使いたい用途には、現状では向きません。 複数のグラフが1つの図に並んでいるものは対象外です。
一致率は「PDFの文字情報と突き合わせられた割合」であり、 文字情報を持たないスキャン文書には別の基準が要ります。

Pricing

You know the price before you commit

A one-off charge per document. No account needed. All output formats (Excel, Word, Markdown) are included, and you can export as many times as you like.

ページ
¥30011〜30ページ
Price tiers
PagesPrice
1〜10¥160
11〜30¥300
31〜60¥550
61〜100¥850
101〜200¥1500
201〜Credits (prepaid)

診断は無料です。料金がかかるのは変換するときだけで、支払い後の追加料金はありません。

How it works

The cause, and the fix

Why it happens

PDFには文字の形(グリフ)と、それがどの文字なのかの対応表が別々に入っています。 この対応表(ToUnicode CMap)が欠けていると、コピーしたときにグリフの番号が そのまま文字として解釈され、まったく別の文字体系に化けます。

印刷やページの見た目には影響しません。だから作成者も気づかないまま公開されています。

How we fix it

対応表が無い以上、テキストとして取り出す方法はありません。画像として読み直すしかありません。

PDFIntact はページを画像として認識し、正しい日本語として書き出します。 表や数式も構造を保ったまま取り出せます。

Stay in touch

Get notified

We'll let you know about accuracy improvements and new output formats. The check itself is free and needs no account.

Used only for these updates. You can unsubscribe at any time.