4 Comments
User's avatar
Matt Gaseltine's avatar

Enjoyed this read. I've noticed myself becoming significantly more optimistic about AI Alignment as a 'Grand Project' actually working in reality. I think this captures a lot of that. I also think there's something to be said for the sheer capabilities of current frontier models joined up with their - as it seems to me - fairly decent understanding and adoption of the 'good' human values.

Sichu Lu's avatar

there are low dimensional ranked structure for those who see high dimensional cathedrals everywhere

Sichu Lu's avatar

every day I wake up and pray alignment is a problem similar enough in computational complexity to the class of human solvable problems

Kurt Pieper's avatar

I think the r=.2 argument with the Likert-Scale may be rescued. A reductio ad absurdum: The correlation between the AFQT (some aptitude test) and the WAIS-IV is r=0.7, so it only explains 49%, a minority, of the variance! Therefore, clearly anyone who talks about "intelligence" being the explanation is just renaming things! It's more the illegible thing that the likert scale is pointing at.

In other words, we may project things we know from legible-space up to illegible-space (the so-called "real world") proportionally.

That being said, something something great men theory of history may be true, idk