Building an open benchmark -
SWE-WebDevBench - for evaluating AI coding platforms like ourselves but also Replit, Lovable etc for real-world webapp development on complete software dev lifecycle.
Created a "lite" version for now, and did a comparative run against 6 major AI coding products. The
Arxiv paper describes the framework.
https://webdevbench.com/