benchmark_id	sku_id	category	product_promise	source_set_version	scene_brief_version	model_route	model_snapshot_or_test_date	candidate_count	accepted_count	identity_drift_rejects	data_drift_rejects	geometry_drift_rejects	appearance_drift_rejects	rights_marking_rejects	claim_drift_rejects	review_minutes	correction_minutes	loaded_labor_rate_usd	generation_cost_usd	external_cost_usd	total_cost_usd	acceptance_rate	cost_per_accepted_asset_usd	decision	evidence_folder	notes
YOUR_BENCHMARK_ID	YOUR_SKU	YOUR_CATEGORY	exact_sku	SOURCE_SET_V1	SCENE_BRIEF_V1	MODEL_ROUTE_1	SNAPSHOT_OR_TEST_DATE	4													CALCULATE	CALCULATE	CALCULATE	NOT_RUN	EVIDENCE_FOLDER_1	Replace placeholders; keep source, brief, crop, size, and candidate count fixed across rows.
YOUR_BENCHMARK_ID	YOUR_SKU	YOUR_CATEGORY	exact_sku	SOURCE_SET_V1	SCENE_BRIEF_V1	MODEL_ROUTE_2	SNAPSHOT_OR_TEST_DATE	4													CALCULATE	CALCULATE	CALCULATE	NOT_RUN	EVIDENCE_FOLDER_2	Record every candidate and every product-truth rejection before judging visual craft.
YOUR_BENCHMARK_ID	YOUR_SKU	YOUR_CATEGORY	exact_sku	SOURCE_SET_V1	SCENE_BRIEF_V1	MODEL_ROUTE_3	SNAPSHOT_OR_TEST_DATE	4													CALCULATE	CALCULATE	CALCULATE	NOT_RUN	EVIDENCE_FOLDER_3	One candidate can trigger multiple rejection classes; do not force failure counts to equal rejected candidates.
YOUR_BENCHMARK_ID	YOUR_SKU	YOUR_CATEGORY	exact_sku	SOURCE_SET_V1	SCENE_BRIEF_V1	MODEL_ROUTE_4	SNAPSHOT_OR_TEST_DATE	4													CALCULATE	CALCULATE	CALCULATE	NOT_RUN	EVIDENCE_FOLDER_4	If accepted_count is zero, mark the route failed instead of reporting a zero cost per accepted asset.
