Apart Research
Subscribe
Sign in
Home
Apart Newsletter
Hackathons
Product development
Organization news
š Versión en espaƱol
Large language models might always be slightly misaligned
Apart Newsletter #28
Apr 25, 2023
Ā
ā¢
Ā
Apart Research
3
Latest
Top
Apart Newsletter #27
ML safety research in interpretability and model sharing, game-playing language models, and critiques of AGI risk research
Apr 18, 2023
Ā
ā¢
Ā
Apart Research
3
Identifying semantic neurons, mechanistic circuits & interpretability web apps
Results from the 4th Alignment Jam
Apr 13, 2023
Ā
ā¢
Ā
CC
2
Ethics or Reward?
This week we take a look at LLMs that need therapists, governance of machine learning hardware, and benchmarks for dangerous behaviour. Read to the endā¦
Apr 9, 2023
Ā
ā¢
Ā
CC
5
Governing AI & Evaluating Danger
We might need to shut it all down, AI governance seems more important than ever and technical research is challenged. Welcome to this week's updateā¦
Apr 3, 2023
What a Week! GPT-4 & Japanese Alignment
What a week. There was already a lot to cover Monday when I came in for work and I was going to do a special feature on the Japan Alignment Conferenceā¦
Mar 15, 2023
Perspectives on AI Safety
This week, we take a look at interpretability used on a Go-playing neural network, glitchy tokens and the opinions and actions of top AI labs andā¦
Mar 6, 2023
Bing Wants to Kill Humanity W07
Welcome to this weekās ML & AI safety update where we look at Bing going bananas, see that certification mechanisms can be exploited and that scalingā¦
Feb 21, 2023
See all
Apart Research
We share updates about the progress of ML safety, run the Apart Lab along with exciting ML safety hackathons.
Subscribe
Apart Research
Subscribe
About
Archive
Sitemap
This site requires JavaScript to run correctly. Please
turn on JavaScript
or unblock scripts