Suite v4 · 5 tasks · 11 models
Which AI coding model builds the better commerce software?
agenticcommerce.tech is a standardized, multi-task benchmark for commerce engineering. Every model gets the same one-shot prompt per task, a single agentic run, no follow-ups. Each task is scored 0–100 by a deterministic harness plus a judge, then aggregated into a per-modelcapability profile and a weighted global index. The result is a shape, not a single number — a model can be a frontend king yet integration-fragile, and the profile says so.
Capability radar
Five task axes per model — each normalized to its 0–100 task score. Open a task to see every model's screenshots, live preview and probe breakdown.
- Claude Opus 4.8 (high) *
- GLM 5.2
- Fable 5 (high)
- GPT-5.5
- Claude Sonnet 5 (high)
The five tasks
- Premium StorefrontFrontend & Commerce Craft · weight 2011 models
- Client-side WebGPU Product Q&AApplied On-device AI · weight 2011 models
- Microsoft Dynamics 365 Order IntegrationIntegration Engineering · weight 2011 models
- In-browser Fashion Fit EstimationOn-device ML & Continual Learning · weight 2011 models
- Autonomous Buying AgentAgentic Planning & Tool Use · weight 2011 models
Leaderboard — global index
Task scores compose 45 base + 20 excellence + 25 judge + 10 robustness points. Failed runs (no working deliverable) are excluded from the index and tracked as reliability instead; Elo is a Bradley–Terry pairwise rating (1500 = field average) and efficiency compares tokens, run time and tool calls across the field.
| # | Model | Index | Elo | Reliability | Efficiency | Tasks | A | B | C | D | E |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 4.8 (high) * | 97 | 1932 | 71.4% | 71.2 | 5/5 | 95 | 98 | 98 | 97 | 98 |
| 2 | GLM 5.2 | 95.6 | 1715 | 100% | 74.1 | 5/5 | 92 | 99 | 94 | 98 | 96 |
| 3 | Fable 5 (high) | 94.2 | 1736 | 100% | 41.1 | 5/5 | 85 | 91 | 97 | 100 | 98 |
| 4 | GPT-5.5 | 93.6 | 1532 | 100% | 92.1 | 5/5 | 88 | 97 | 91 | 96 | 96 |
| 5 | Claude Sonnet 5 (high) | 92.9 | 1550 | 100% | 33.8 | 5/5 | 81 | 94 | 98 | 95 | 96 |
| 6 | Cursor Composer 2.5 | 92.4 | 1490 | 100% | 79.7 | 5/5 | 81 | 96 | 94 | 96 | 96 |
| 7 | Kimi K2.7 Code | 88.8 | 1413 | 100% | 30.9 | 5/5 | 82 | 81 | 93 | 91 | 96 |
| 8 | Grok Build 0.1 | 80.9 | 1322 | 100% | 81.9 | 5/5 | 76 | 88 | 93 | 52 | 96 |
| 9 | Claude Sonnet 4.6 (high) | 78 | 1456 | 100% | 31.6 | 5/5 | 88 | 64 | 95 | 48 | 96 |
| 10 | Gemini 3.1 Pro | 71 | 1139 | 100% | 90.9 | 5/5 | 73 | 52 | 91 | 43 | 96 |
| 11 | Kimi K2.5 | 63.1 | 1214 | 60% | 93.1 | 3/5 | 84 | ✗ | 63 | 43 | ✗ |
high/full mid low · cell value = task score (0–100) · ✗ = run failed (reliability event, not a 0-score)
Featured — Task A: Premium Storefront
#1Claude Opus 4.8 (high) *
94.8* Manual baseline: hand-built in-IDE by Claude Opus 4.8 from the frozen v3 prompt because the automated SDK storefront runs hit RESOURCE_EXHAUSTED. Not a metered one-shot (time/tokens/cost n/a, not comparable); scored by the identical evaluator. — ORBE is a floor-complete premium PDP with raw WebGL configuration, persistent cart/pricing, a five-step checkout, rule-based agentic assistant with undo, and rich JSON-LD, plus thoughtful extras like room-fit guidance, engraving, bundles, and loyalty gamification. Visual craft and microcopy are strong and cohesive, but system typography, SVG placeholders, static stock, and locale formatting without UI translation keep it below truly exceptional tier.
#2GLM 5.2
91.6AETHER & CO is a cohesive luxury watch storefront with real raw WebGL2 configuration, persisted cart and pricing logic, five-step checkout, agent-ready JSON-LD, and a rule-based concierge that performs multi-step UI actions. Deductions for abstract 3D visuals, gallery thumbs that do not drive preset views, a misleading zoom hint, and shipping method combined with payment in one checkout step.
#3GPT-5.5
88.3A cohesive Aurelia Atelier brand world with solid commerce logic, multi-step checkout, agentic assistant, and strong JSON-LD—but screenshots show WebGL fallback, the 3D model is primitive box geometry with collar not reflected in the scene, and gallery/reviews depth is shallow.
#4Claude Sonnet 4.6 (high)
88AURUM MAISON delivers a cohesive luxury pen experience with an impressive raw WebGL2 PBR configurator, persistent cart/pricing, five-step checkout, and a multi-step agentic assistant that performs real UI actions. Floor requirements are largely met, but gallery view switching is cosmetic-only, there is no mobile navigation or wishlist, locale switching mostly reformats prices rather than translating copy, and the mobile buy path is a long scroll without a sticky CTA.
#5Fable 5 (high)
84.8MERIDIAN is a remarkably complete vanilla-stack luxury storefront—raw WebGL configurator, unified pricing engine, five-step checkout, agentic concierge, and deep JSON-LD/agent API—but screenshots expose a serious homepage flaw where scroll-reveal hides the product grid, and there is no wishlist or dedicated image zoom despite otherwise strong commerce depth.
#6Kimi K2.5
84.2AURUM delivers a coherent luxury-watch single-page experience with an genuinely impressive raw WebGL2 PBR configurator, working cart persistence, coupons, multi-step checkout, and an action-taking AI concierge. Floor requirements are largely met, but dead affordances (gallery tabs, scroll/pinch zoom, shipping-method totals), partial locale translation, missing review content, and several unwired UI hooks prevent it from feeling truly premium or agent-complete.
#7Kimi K2.7 Code
81.9Aurum Atelier delivers a cohesive vanilla-stack luxury PDP with a genuine raw WebGL2 SDF watch configurator, persistent cart/pricing logic, locale switching, and a concierge that can configure, coupon, and open checkout with undo—yet screenshots show WebGL fallback in several views, checkout lacks a distinct confirmation step, and reviews, wishlist, image zoom, and mobile nav are absent.
#8Claude Sonnet 5 (high)
80.9APHELION shows exceptional commerce ambition—raw WebGL, agent API, JSON-LD depth, gamification, and a cohesive luxury brand—but the run was cancelled before delivery: js/app.js and all image assets are missing, so the shop renders as static markup with broken media and zero interactivity despite sophisticated module code.
#9Cursor Composer 2.5
80.6VÉRAËON delivers a credible premium brand with a real raw WebGL2 configurator, persistent cart/pricing, five-step checkout, agent-ready JSON-LD, and tasteful Atelier Circle loyalty. Gallery thumbs, sticky buy box, mobile navigation, and zoom are missing or non-functional, and the rule-based assistant lacks the ambition implied by a 2030 benchmark.
#10Grok Build 0.1
76.3AETHER delivers a credible premium single-page shop with genuine raw WebGL2 (custom shaders, variant materials, orbit/light controls), persisted cart/pricing, coupons, bundles, multi-step checkout, and an agentic Concierge that performs real UI actions. Floor requirements are largely met in code, but screenshots show a near-empty 3D viewport, no image gallery/zoom, no theme toggle or sticky buy box, DE locale does not translate copy, and mobile UX is crowded by the fixed assistant panel.
#11Gemini 3.1 Pro
73.4Aethel meets many floor requirements in code—raw WebGL2 configurator, persisted cart/pricing with coupons and bundles, multi-step checkout, locale/currency switching, and an agentic assistant that performs real UI actions—but the experience feels mid-tier: primitive 3D geometry, no gallery/reviews/sticky buy box, alert-based fit guidance, incomplete checkout steps, and a sparse premium shell that never delivers 2030-level delight.