Tutorials
Claude Opus 5.5: Turn Product PDFs Into Checkable Spec Tables
Use Claude Opus 5.5 to extract a source-linked specification table before asking it to write product comparisons.

Use Claude Opus 5.5 to extract a source-linked specification table before asking it to write product comparisons. Preserve the original value, unit, model and page beside each normalized value. An attractive comparison grid is not useful if a shipping weight becomes a product weight or an empty warranty cell becomes zero years.
This guide provides a proposed extraction workflow and invented examples, not a benchmark of Opus accuracy. Claude's PDF documentation describes reading text and visual content from documents. It does not make an extracted specification automatically reliable.
Start with one representative page
Choose a page containing a merged header, a footnote or more than one product variant. Supply the original PDF and identify both the PDF viewer page number and the printed page label when they differ. Keep the document filename and revision date in the extraction record.
Ask for the column hierarchy first. A header such as “Model A / compact / extended” may apply to several subcolumns. A plain-text dump can lose that relationship. The pypdf extraction documentation explains why PDF layout and semantic table structure are not equivalent. Inspect the page image when a value appears under the wrong heading.
A Claude Code issue about engineering tables reports this difficulty with an earlier Sonnet workflow. It is a user report, not an Opus 5.5 error-rate measurement, and its claimed speedup is not adopted here.
Keep an evidence column
| Field | Hypothetical source | Normalized value | Required check |
|---|---|---|---|
| Product mass | 0.75 kg, page 3 | 750 g | Net mass, not boxed mass |
| Width | 12 cm, page 4 | 120 mm | Width axis, not depth |
| Warranty | Blank cell, page 6 | Unknown | Do not interpret as no warranty |
| Runtime | Up to 8 h, footnote 2 | Up to 8 h | Keep operating condition |
The arithmetic is straightforward: 0.75 × 1,000 = 750 and 12 × 10 = 120. The harder question is whether the numbers describe the same thing. Keep gross/net, typical/maximum and included/optional distinctions as separate fields. Do not turn a rounded source measurement into extra apparent precision.
For an ambiguous cell, record the original text and “needs review.” Guessing a likely specification makes the comparison harder to audit later. Missing, illegible, conflicting and not applicable are different states; a single empty output field conceals those differences.
Copy this extraction prompt
Read the attached product PDF. First identify the document title, revision date, product variants and the table's column hierarchy on viewer page 3.
Extract only the requested fields: model, net mass, width, depth, height, runtime and warranty.
For every field return: model, original text, original unit, normalized value, normalized unit, PDF viewer page, printed page label, footnote and status.
Use g for mass and mm for dimensions. Show conversion arithmetic. Preserve “up to,” “typical,” test conditions and stated precision.
Do not infer a missing value from another variant. Use distinct statuses for missing, illegible, conflicting and not applicable. Do not replace missing values with zero.
Separate extraction from comparison. Return unresolved cells first, then the table. Do not write advertising claims yet.
Resolve conflicts before making claims
Suppose the overview says 750 g and a shipping page says 1.1 kg. Those values may both be correct for different scopes. Ask Claude to show the heading and neighboring text for each; do not choose whichever makes the product look lighter.
If two revisions disagree, select the approved current document through the product owner or existing release record. Record the older value as superseded rather than silently combining specifications from both editions. Once the pilot page is correct, process the remaining pages using the same field schema.
Review every field that will appear in the final ad against its source page, including footnotes. A comparison can claim “250 g lighter” only after the two approved net masses have matching scope; it cannot claim “better portability” solely from that subtraction without an explicit editorial interpretation.
Take approved facts into the visual brief
Create a short approved-facts block containing model identifiers, values, units and required qualifications. Bring that block and the correct product image into Panelly Studio to build a complete comparison image. Keep unknown cells out of the visual claim, or label them explicitly when they matter to the comparison.
This handoff does not imply Panelly reads PDFs through Opus. It separates document extraction from visual production. See product-variant mixups for keeping images bound to the right SKU, and feature callout images for deciding which verified facts deserve visual emphasis.
Common questions
Is OCR text enough for a scanned table?
It may recover characters while losing their relationships. Inspect the page and compare the extracted row and column labels before accepting values.
Should I upload every document at once?
Start with a representative page and a fixed schema. Scale after that page passes review; a larger batch does not repair a misunderstood header.


