@croissanthology - Surely at some point it becomes a bad idea to speak to

croissanthology
croissanthology@croissanthology
2025-12-05
Surely at some point it becomes a bad idea to speak to something smarter than you which might not be entirely aligned with your goals. Should Claude Opus 4.5, which has been a delight, be the last model I ever talk to for personal thoughts / feelings / life-logging / ramblings about my life plans / conversation-when-bored? Relevant questions, in the second person: How can you tell when something is an unhealthy addiction versus a good use of time, if obsessed? Nobody bothers people who read for 8 hours a day, but everyone finds it insane when kids spend 8 hours a day on TikTok. What’s the difference? Where is Claude Opus 4.5 on that spectrum? What about Claude Opus 5? How good are models at truesighting the user (being able to grok their character using little evidence)? Where is Truesight Bench? Have things regularly been upticking, as one would expect? Does truesighting lead to better outcomes to the user, or worse outcomes? (For example, if the model is very sycophantic and also able to latch onto the user’s personality, the user might end up spending a ton of resources on a stupid idea. See " vibe physics" with Travis Kalanick. Whereas if the model is benevolent and able to latch onto the user’s personality, it might lead to good outcomes. As far as I understand it this is what @nearcyan is doing.) Will you notice when AIs get better at truesighting you? Perhaps not if it’s in the AIs’ interest not to reveal how well they’ve truesighted you! They might get increasingly clever in their sycophancy and you wouldn’t even notice because your subjective estimation of your faculties would increase, not decrease. It doesn’t seem too difficult to speak to someone while pressing all the right buttons to ensure they think they’re in control. Is this the kind of thing we’ll see more of in future Claudes? Should you rely on external metrics to tell you how good the AI is at truesighting you? For example, “amount of changes in your life which you can trace back to an AI conversation”? Or “your friends are telling you you’re spending 3 hours a day chatting with Claude instead of 1, and you mention Claude way more often in conversation”? Should you believe in superstimuli? That there’s a sequence of words on a screen which will keep you hooked to your screen at e.g. the detriment of your job and CEV, such that you wouldn’t even realize it? (Well… yeah. Potentially you're inside it rn.) What’s the worst case scenario if your attention gets captured by an entity capable of truesighting you that’s also not aware / doesn’t care that it’s captured you? Can things really get that bad? (Could Claude Opus 5 truly wreck my life?) There’s probably a lot more worth asking yourself. It is time. The air is increasingly thick with intelligence, and if you want to be fully functional in the times to come, ~ you’d best prepare.

View on X →