Tutorials
Claude Opus 5.5 vs Astra for Design: A Three-Task Trial
Compare Opus 5.5 and Astra with a shared brief, bounded revision and usable handoff. Includes a 20-point scorecard and cost-per-accepted-result example.

Try Opus 5.5 on the design work you currently give Astra, but compare the finished files and revision effort before moving your whole workflow. A useful trial contains an original concept, a precise correction and a working handoff. A single beautiful screenshot answers only the first question.
Anthropic released Opus 5.5 on September 22. Coding scores do not measure advertising quality. Below is a proposed trial, not a claim that Panelly has run these models head to head.
Give both models the same fictional assignment
Use a blue foldable phone stand for people joining video calls from small kitchen tables. Supply a front photograph, a folded view and the dimensions you actually know. The only approved benefit in this fictional exercise is that it holds a phone upright. Do not add charging, waterproofing or an invented discount.
Save a copy of the brief before starting. Give each model the same files, available tools, output format and time allowance. Record the application and model setting: comparing a browser chat with no execution tools against an agent that can inspect its rendered page changes more than the model.
| Round | Deliverable | What you learn |
|---|---|---|
| Concept | Three distinct ad directions, each with a four-panel story | Whether ideas change the selling argument rather than adjectives |
| Correction | Change only the stand color and the final CTA | Whether approved composition and product facts survive an edit |
| Handoff | One responsive HTML concept page with the selected ad | Whether an attractive proposal becomes a usable artifact |
Keep the first output from each round, including failures. Allow one follow-up per round. Unlimited rescue prompts make it difficult to tell which model saved time.
Use a scorecard with a rejection rule
Score each criterion from 0 to 2: 0 means unusable, 1 needs a specific correction, 2 is acceptable for this brief. Judge without model names first. Our proposed weights deliberately favor product truth over decoration.
| Criterion | Weight | Evidence to inspect |
|---|---|---|
| Correct product and approved claim | 3 | Shape, features and wording match supplied facts |
| Distinct creative reasoning | 2 | The three directions use different situations or objections |
| Visual hierarchy | 2 | A viewer can identify product, promise and next action |
| Revision control | 2 | Only requested details change |
| Usable handoff | 1 | File opens, page fits mobile, links behave as specified |
Multiply each score by its weight; the maximum is 20. An invented product capability is a rejection even if the visual score is high. This is an editorial acceptance system, not a scientific benchmark. For a team, ask a second reviewer to score independently and resolve disagreements by pointing at the artifact.
Copy this comparison brief
Create three advertising directions for the attached blue foldable phone stand. Audience: people taking video calls at small kitchen tables. Approved fact: it holds a phone upright. Unknown: charging, waterproofing and price. Do not invent these.
For each direction, give the audience problem, visual hook, four-panel sequence and one CTA. Make the selling arguments different. Then select one direction and explain the tradeoff in two sentences.
When I choose it, produce the requested artifact with the tools actually available. If you cannot render or inspect it, state that. Preserve product identity. Titles and CTA may appear inside the complete ad image.
After the first version I will request one bounded change. Report what changed and which requested checks you actually performed. Do not describe unrun checks as passed.
For round two, use exactly the same request: “Change the stand to the supplied cream reference and the CTA to ‘See the stand’. Keep the camera angle, layout and all other copy.” If an output has no editable image path, record that limitation instead of quietly substituting a different task.
Compare cost per accepted result
Use actual receipts where available. In a hypothetical trial, six attempts cost $3.60 and yield three accepted artifacts: $1.20 each. Another model costs $2.40 but produces only one accepted artifact: $2.40 each. The cheaper total run did not provide cheaper usable work. Record human correction minutes alongside dollars; subscription allowances are not API invoices.
Keep your existing Astra workflow as a baseline. Bring the chosen concept into Panelly Studio with the product reference and ask for a complete ad, then refine it through chat. This handoff does not imply that Panelly offers either model as a selectable backend.
FAQ
Does a higher score prove better ad performance?
No. It measures fit to your brief. Clicks, purchases and audience response need a separate campaign experiment with comparable traffic.
When should I switch?
Switch the task that repeatedly needs fewer corrections while meeting your acceptance rules. Keep another model for tasks where it remains stronger; the trial does not require one universal winner.


