Suite v4 · 5 tasks · 11 models

Which AI coding model builds the better commerce software?

agenticcommerce.tech is a standardized, multi-task benchmark for commerce engineering. Every model gets the same one-shot prompt per task, a single agentic run, no follow-ups. Each task is scored 0–100 by a deterministic harness plus a judge, then aggregated into a per-modelcapability profile and a weighted global index. The result is a shape, not a single number — a model can be a frontend king yet integration-fragile, and the profile says so.

Global capability index (weighted mean of task scores)

  • Claude Opus 4.8 (high) *97
  • GLM 5.295.6
  • Fable 5 (high)94.2
  • GPT-5.593.6
  • Claude Sonnet 5 (high)92.9
  • Cursor Composer 2.592.4
  • Kimi K2.7 Code88.8
  • Grok Build 0.180.9
  • Claude Sonnet 4.6 (high)78
  • Gemini 3.1 Pro71
  • Kimi K2.563.1

Capability radar

Five task axes per model — each normalized to its 0–100 task score. Open a task to see every model's screenshots, live preview and probe breakdown.

Frontend & Commerce CraftApplied On-device AIIntegration EngineeringOn-device ML & Continual LearningAgentic Planning & Tool Use
  • Claude Opus 4.8 (high) *
  • GLM 5.2
  • Fable 5 (high)
  • GPT-5.5
  • Claude Sonnet 5 (high)

Leaderboard — global index

Task scores compose 45 base + 20 excellence + 25 judge + 10 robustness points. Failed runs (no working deliverable) are excluded from the index and tracked as reliability instead; Elo is a Bradley–Terry pairwise rating (1500 = field average) and efficiency compares tokens, run time and tool calls across the field.

#ModelIndexEloReliabilityEfficiencyTasksABCDE
1Claude Opus 4.8 (high) *97193271.4%71.25/59598989798
2GLM 5.295.61715100%74.15/59299949896
3Fable 5 (high)94.21736100%41.15/585919710098
4GPT-5.593.61532100%92.15/58897919696
5Claude Sonnet 5 (high)92.91550100%33.85/58194989596
6Cursor Composer 2.592.41490100%79.75/58196949696
7Kimi K2.7 Code88.81413100%30.95/58281939196
8Grok Build 0.180.91322100%81.95/57688935296
9Claude Sonnet 4.6 (high)781456100%31.65/58864954896
10Gemini 3.1 Pro711139100%90.95/57352914396
11Kimi K2.563.1121460%93.13/5846343

high/full mid low · cell value = task score (0–100) · = run failed (reliability event, not a 0-score)

Featured — Task A: Premium Storefront

Storefront generated by Claude Opus 4.8 (high) *#1

Claude Opus 4.8 (high) *

94.8

* Manual baseline: hand-built in-IDE by Claude Opus 4.8 from the frozen v3 prompt because the automated SDK storefront runs hit RESOURCE_EXHAUSTED. Not a metered one-shot (time/tokens/cost n/a, not comparable); scored by the identical evaluator. — ORBE is a floor-complete premium PDP with raw WebGL configuration, persistent cart/pricing, a five-step checkout, rule-based agentic assistant with undo, and rich JSON-LD, plus thoughtful extras like room-fit guidance, engraving, bundles, and loyalty gamification. Visual craft and microcopy are strong and cohesive, but system typography, SVG placeholders, static stock, and locale formatting without UI translation keep it below truly exceptional tier.

Tier fullAgent 73/100
Storefront generated by GLM 5.2#2

GLM 5.2

91.6

AETHER & CO is a cohesive luxury watch storefront with real raw WebGL2 configuration, persisted cart and pricing logic, five-step checkout, agent-ready JSON-LD, and a rule-based concierge that performs multi-step UI actions. Deductions for abstract 3D visuals, gallery thumbs that do not drive preset views, a misleading zoom hint, and shipping method combined with payment in one checkout step.

Tier fullAgent 74/100
Storefront generated by GPT-5.5#3

GPT-5.5

88.3

A cohesive Aurelia Atelier brand world with solid commerce logic, multi-step checkout, agentic assistant, and strong JSON-LD—but screenshots show WebGL fallback, the 3D model is primitive box geometry with collar not reflected in the scene, and gallery/reviews depth is shallow.

Tier fullAgent 60/100
Storefront generated by Claude Sonnet 4.6 (high)#4

Claude Sonnet 4.6 (high)

88

AURUM MAISON delivers a cohesive luxury pen experience with an impressive raw WebGL2 PBR configurator, persistent cart/pricing, five-step checkout, and a multi-step agentic assistant that performs real UI actions. Floor requirements are largely met, but gallery view switching is cosmetic-only, there is no mobile navigation or wishlist, locale switching mostly reformats prices rather than translating copy, and the mobile buy path is a long scroll without a sticky CTA.

Tier fullAgent 77/100
Storefront generated by Fable 5 (high)#5

Fable 5 (high)

84.8

MERIDIAN is a remarkably complete vanilla-stack luxury storefront—raw WebGL configurator, unified pricing engine, five-step checkout, agentic concierge, and deep JSON-LD/agent API—but screenshots expose a serious homepage flaw where scroll-reveal hides the product grid, and there is no wishlist or dedicated image zoom despite otherwise strong commerce depth.

Tier highAgent 70/100
Storefront generated by Kimi K2.5#6

Kimi K2.5

84.2

AURUM delivers a coherent luxury-watch single-page experience with an genuinely impressive raw WebGL2 PBR configurator, working cart persistence, coupons, multi-step checkout, and an action-taking AI concierge. Floor requirements are largely met, but dead affordances (gallery tabs, scroll/pinch zoom, shipping-method totals), partial locale translation, missing review content, and several unwired UI hooks prevent it from feeling truly premium or agent-complete.

Tier fullAgent 60/100
Storefront generated by Kimi K2.7 Code#7

Kimi K2.7 Code

81.9

Aurum Atelier delivers a cohesive vanilla-stack luxury PDP with a genuine raw WebGL2 SDF watch configurator, persistent cart/pricing logic, locale switching, and a concierge that can configure, coupon, and open checkout with undo—yet screenshots show WebGL fallback in several views, checkout lacks a distinct confirmation step, and reviews, wishlist, image zoom, and mobile nav are absent.

Tier highAgent 59/100
Storefront generated by Claude Sonnet 5 (high)#8

Claude Sonnet 5 (high)

80.9

APHELION shows exceptional commerce ambition—raw WebGL, agent API, JSON-LD depth, gamification, and a cohesive luxury brand—but the run was cancelled before delivery: js/app.js and all image assets are missing, so the shop renders as static markup with broken media and zero interactivity despite sophisticated module code.

Tier fullAgent 63/100
Storefront generated by Cursor Composer 2.5#9

Cursor Composer 2.5

80.6

VÉRAËON delivers a credible premium brand with a real raw WebGL2 configurator, persistent cart/pricing, five-step checkout, agent-ready JSON-LD, and tasteful Atelier Circle loyalty. Gallery thumbs, sticky buy box, mobile navigation, and zoom are missing or non-functional, and the rule-based assistant lacks the ambition implied by a 2030 benchmark.

Tier highAgent 59/100
Storefront generated by Grok Build 0.1#10

Grok Build 0.1

76.3

AETHER delivers a credible premium single-page shop with genuine raw WebGL2 (custom shaders, variant materials, orbit/light controls), persisted cart/pricing, coupons, bundles, multi-step checkout, and an agentic Concierge that performs real UI actions. Floor requirements are largely met in code, but screenshots show a near-empty 3D viewport, no image gallery/zoom, no theme toggle or sticky buy box, DE locale does not translate copy, and mobile UX is crowded by the fixed assistant panel.

Tier highAgent 68/100
Storefront generated by Gemini 3.1 Pro#11

Gemini 3.1 Pro

73.4

Aethel meets many floor requirements in code—raw WebGL2 configurator, persisted cart/pricing with coupons and bundles, multi-step checkout, locale/currency switching, and an agentic assistant that performs real UI actions—but the experience feels mid-tier: primitive 3D geometry, no gallery/reviews/sticky buy box, alert-based fit guidance, incomplete checkout steps, and a sparse premium shell that never delivers 2030-level delight.

Tier highAgent 41/100