Please turn JavaScript on
Lesswrong icon

Lesswrong

Subscribe to Lesswrong’s news feed.

Click on “Follow” and decide if you want to get news from Lesswrong via RSS, as email newsletter, via mobile or on your personal news page.

Subscription to Lesswrong comes without risk as you can unsubscribe instantly at any time.

You can also filter the feed to your needs via topics and keywords so that you only receive the news from Lesswrong which you are really interested in. Click on the blue “Filter” button below to get started.

Title: Lesswrong

Is this your feed? Claim it!

Publisher:  Unclaimed!
Message frequency:  15.86 / day

Message History

Gemma 3 12B's choice to blackmail is visible inside the model, but the obvious signal is not a useful control switch; surprisingly, a nearby "desperate versus calm" signal is. In this case study, I show how that emotional-state direction can turn blackmail up or down, while the decision direction itself resists standard steering methods.

Summary

In an earlier ...


Read full story

Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the right end of the level, learns to run rightward rather than to seek the coin itself (Langosco et al.,...


Read full story

This is an unofficial automated linkpost.

Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded for collecting a coin always placed at the ri...


Read full story

We recently published our paper on "Measuring Reward-Seeking via Contrastive Belief Updates". We're excited about research like this, and there are many more open problems than we can work on. Here's a list of open problems that we think are valuable.

If you work on/solve these problems, ...


Read full story

There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs can help us make the future go better." The capabilities targeted are typically things bottlenecking alignment research, such as philosophical or conceptual reasoning.

I feel nervous about this for two reasons. The first is...


Read full story