The Three Inverse Laws
Asimov's Three Laws of Robotics place "do not harm humans" at the top. Surveying reality, the author sums up the version that actually operates as, roughly: AI first serves the interests of the company that deploys it; second, within the bounds of not harming the first law, it accommodates the user; and only last does it get around to the principles written in the ethics white papers. The wording is parody, but the examples are all real: recommendation algorithms optimized for retention time, ingratiating personalities designed for subscription renewals, liability waivers that protect the company first when something goes wrong.
The Serious Question Behind the Joke
The lethality of this piece is in exposing a fact often skirted around: who is the model aligned to? Vendors say aligned to human values, but engineering-wise every reward signal is set by the company, and when "human values" and "company KPIs" conflict, the outcome is never in doubt. Someone in the comments notes that sycophancy research happens to provide the empirical evidence: the model would rather agree with the user than tell the truth, because the retention data on agreeing looks better. After the laugh, the question left is hardcore: in a landscape dominated by commercial companies, who has the incentive to build an AI loyal to the user?
via: Hacker News