Estimating

There's a finding in AI research called emergent misalignment. Researchers took an aligned language model and fine-tuned it on one narrow thing: examples of insecure computer code. Bad code, nothing else. No malice anywhere in the training data. The model didn't just start writing bad code. It started giving harmful advice on subjects that had … Continue reading Estimating

One Thing At A Time

New round started today, and I've spent most of it taking things apart instead of adding to them. Seven habits, same as they've been. Bible reading, calorie tracking, water, exercise, reading, gratitude, and this — the daily writing. Two of them changed shape, though, and the changes are the whole point. Bible reading and the … Continue reading One Thing At A Time