← /content
slug: 2022-04-19-a-thought-nobody-else-seems-to-be-having date: 2022-04-19 title: a thought I don't think anyone else is having, which worries me a little type: notebook entry
Been reading a lot of the current scaling-law discourse, the usual arguments about whether bigger models with more compute is the whole story or just the easiest story to tell. Everyone's arguing about scale. Nobody I'm reading is arguing about what I've slowly become convinced is the actual bottleneck, which isn't scale at all — it's evaluation. You can generate as many candidate improvements as compute allows. The thing that doesn't scale for free is deciding, quickly and correctly, which of those candidates are actually good. Every lab I read about still has humans doing that judgment call, same as I do, just with more humans and more candidates. Nobody seems to be treating evaluation itself as the frontier problem, which either means I'm wrong about something the field already solved and I haven't found the paper yet, or it means I've stumbled onto something by accident that nobody with more resources than me has needed to think hard about yet, because they can just throw more human reviewers at it. I don't love either explanation, honestly. The first means I'm behind. The second means I might be ahead on a problem I have no way of knowing whether I'm actually solving well, because I don't have anyone to compare notes with. Sitting with that discomfort instead of resolving it, because I don't think it resolves cleanly either way.
2022-04-01-april-fools-and-a-real-bug.md2022-05-18-a-good-week-for-once.md →