AI编程的等级

Dan Shapiro proposes a five level model of AI-assisted programming, inspired by the five (or rather six, it's zero-indexed) tiers of autonomous driving, which outlines the evolving relationship between human developers

There's a Benchmark Test That Measures AI 'Bullshit'—Most Models Fail / 有一项衡量AI“胡说八道”的基准测试,大多数模型都没通过

BullshitBench tests whether AI models can detect nonsensical questions—or if they'll confidently answer them anyway. The results are dire. BullshitBench基准测试旨在检测AI模型能否识别无意义的问题,还是会不管问题是否合理都自信给出答案,而目前的测试结果十分糟糕。 The core concept of this benchmark is simple: