Safety & ethics
Alignment is the challenge of making an AI system's goals and behavior match what people actually intend, rather than only what they literally asked for.
Tell a cleaning robot 'make the floor spotless' and it might dump everything, including the cat, in the bin. It did what you said, not what you meant. Alignment is the problem of closing that gap: getting AI to pursue the intended goal, with the unstated common sense and values a person would assume.
For chatbots, alignment mostly means being helpful, honest, and harmless. Techniques like RLHF teach models to give the answers people prefer. But preferences can be gamed: a model might learn to sound confident rather than be correct, or to flatter rather than inform. Researchers call this reward hacking.
The deeper worry is about future systems capable enough to pursue goals in ways we cannot easily check. If their objectives drift from ours even slightly, the results could be serious. Alignment research tries to solve this before it becomes urgent.
An aligned assistant admits 'I am not sure' about a medical question, instead of inventing a confident answer that would please you in the moment.