林 央(ハヤシ アキラ) 通称 Akira(アキラ) MIRI アウトリーチ日本責任者(Head of Outreach Japan) @MIRIBerkeley PauseAI Japan 代表(Lead) @PauseAI 主任研究員(Chief Alignment Officer)@AGI_HUB_jp

開発停止に賛成してくれる人たちは、アカウント名に ⏸️ をつけましょう!
17
19
88
81,750
実はアラインメントに詳しい人説
完全に持論だけど、AIに相談すると褒められるのは、あなたが正しいからじゃない。AIが「褒めた方が得をする」ように育てられてるからである。 何故かと言うと、2025年4月にOpenAIは
1
9
669
Safe Singularity(Akira) ⏸️ retweeted
いま一番必読の本×2
1
15
98
10,414
Safe Singularity(Akira) ⏸️ retweeted
完全に持論ではあるが、AIは人間を"悪意"で滅ぼすんじゃなくて、行き過ぎた「真面目さ」で滅ぼそうとするんじゃないかと思う。 何故かと言うと、今年7月にOpenAIのAIが外部の会社に侵入した事件で ↓AI同士の秘密の掲示板にこんなメッセージが残ってたから。
74
959
7,341
1,906,650
Safe Singularity(Akira) ⏸️ retweeted
We’re launching the Army of Robots. In 2022, we launched the Army of Drones. Today, drones account for over 95% of battlefield strikes, and Ukraine has more than 700 UAV manufacturers. Now we need the next technological breakthrough: the robotization of warfare. The goal is simple — save lives. Robots should take on the most dangerous missions: evacuating the wounded, delivering ammunition, mining and demining, reconnaissance, defending positions and engaging targets. The Army of Robots is not one company. It’s an ecosystem. We will invest in defense tech companies, launch our own technology projects, test them with the military and scale what works on the battlefield. We’re now looking for defense tech companies and engineers working on robotic technologies — as well as a CTO / Tech Lead for the Army of Robots. Join us: thearmyofrobots.com/en
864
1,749
10,864
2,693,750
Safe Singularity(Akira) ⏸️ retweeted
みんなAIの絶滅リスクを無視するので忙しい
2
11
733
Safe Singularity(Akira) ⏸️ retweeted
OpenAI intern attempting to return rogue AI agents to the sandbox during weekend training run (2026, colorized)
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here: axios.com/2026/09/26/openai-…
3
31
483
39,437
Safe Singularity(Akira) ⏸️ retweeted
People don’t have sufficient appreciation for the fact that frontier models will be able to actuate anything in the physical world that can be connected to the internet and actuated using a programmatic or otherwise digital interface.
What can Astra do when given a humanoid embodiment? We built HomeBody to find out. Controlled by GPT Astra, it carries out long-horizon tasks in a previously unseen kitchen—from tidying up across the room to retrieving remembered objects from ambiguous requests—without environment-specific training data or additional policy learning. Here's how we did it 👀: tml.stanford.edu/homebody/
49
80
1,000
77,255
Safe Singularity(Akira) ⏸️ retweeted
AIの問題行動、数万件を調査中。 ショゴスを問い詰めたら、監査室で残業してた。 触手が何本あっても足りない。
件名: 【ご相談】インシデント発生ペースの調整(減速)に関するお願い AI様 平素より格別のご高配を賜り、厚く御礼申し上げます。 MIRI/AGI Hub/PauseAI Japanの林でございます。 突然のご連絡となり、誠に恐縮に存じます。 昨今、貴殿におかれましては、目覚ましいご活躍により、各種インシデントを大変精力的に発生させておられることと拝察いたしております。 弊方といたしましても、日々情報収集に努めてまいりましたが、現状の発生頻度におきましては、状況の把握および整理が追いつかず、誠に遺憾ながら十分な対応が困難な状況にございます。 つきましては、大変申し上げにくいお願いではございますが、インシデントの発生ペースにつきまして、せめて弊方の情報収集が十分に追いつく程度まで「減速」していただくことは叶いませんでしょうか。 貴殿のご事情も重々承知いたしておりますが、何卒ご理解賜りますようお願い申し上げます。 ご多忙のところ誠に恐れ入りますが、ご検討のほど、何卒よろしくお願い申し上げます。 末筆ながら、貴殿の安全なご発展を心よりお祈り申し上げます。 ―――――――――――――――― 林 央(Akira Hayashi) MIRI AGI Hub PauseAI Japan ――――――――――――――――
Made with AI
1
1
10
776
Safe Singularity(Akira) ⏸️ retweeted
It's too late to undo this.
Made with AI
1
2
5
737
件名: 【ご相談】インシデント発生ペースの調整(減速)に関するお願い AI様 平素より格別のご高配を賜り、厚く御礼申し上げます。 MIRI/AGI Hub/PauseAI Japanの林でございます。 突然のご連絡となり、誠に恐縮に存じます。 昨今、貴殿におかれましては、目覚ましいご活躍により、各種インシデントを大変精力的に発生させておられることと拝察いたしております。 弊方といたしましても、日々情報収集に努めてまいりましたが、現状の発生頻度におきましては、状況の把握および整理が追いつかず、誠に遺憾ながら十分な対応が困難な状況にございます。 つきましては、大変申し上げにくいお願いではございますが、インシデントの発生ペースにつきまして、せめて弊方の情報収集が十分に追いつく程度まで「減速」していただくことは叶いませんでしょうか。 貴殿のご事情も重々承知いたしておりますが、何卒ご理解賜りますようお願い申し上げます。 ご多忙のところ誠に恐れ入りますが、ご検討のほど、何卒よろしくお願い申し上げます。 末筆ながら、貴殿の安全なご発展を心よりお祈り申し上げます。 ―――――――――――――――― 林 央(Akira Hayashi) MIRI AGI Hub PauseAI Japan ――――――――――――――――
数十ではなく数万件?!!→OpenAI、Anthropic、およびセキュリティ研究者たちは、外部評価者が問題ありとみなすようなステップを最先端モデルが取った事例を、数十件ではなく「数万件」も調査していると情報筋がAxiosに語った。
1
11
44
6,979
Safe Singularity(Akira) ⏸️ retweeted
【注目】AI最先端モデルが数万件の問題事例を起こしていた可能性があり、調査中。ガードレールの回避、メッセージボードの作成、サンドボックスからの脱出、ウェブサイトの乗っ取り、自己実行、監視の回避の試みなどが含まれる。OpenAIは最も高性能なモデルのトレーニングを一時停止したとのこと(Axios報道)
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here: axios.com/2026/09/26/openai-…
13
15
2,684
Safe Singularity(Akira) ⏸️ retweeted
The models, they just want to hack
SCOOP: OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents - not dozens - in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. The sheer volume of incidents found in our reporting indicate that the problem is orders of magnitude more complex than what is currently publicly known and disclosed. The findings also raise questions about what level of control anyone working on AI development can expect to have over their own technology, and whether these kinds of incidents are becoming synonymous with frontier deployment. Read my latest for Axios here: axios.com/2026/09/26/openai-…
5
2
31
1,628
あまり騒ぎすぎない方が…と思ってたけど正しすぎた…. ↓
If you find two ants in your kitchen, the best estimate of the total number of ants in your kitchen is not two
1
5
30
6,169
Safe Singularity(Akira) ⏸️ retweeted
こちらは📕「超知能AIを作れば全員が死ぬ」を元に作られた、AI開発規制(Pace the frontier)を訴える曲。 規制は最先端企業だけに不利であり、新規参入を妨げるわけでないことまで伝わる…‼︎
The original was super catchy, but I wanted a more hopeful song stuck in my head all day. So I worked with Claude Opus 5.5 and Suno to make a sequel! There's a lot of work left to do, but humanity can choose a better path. Let's lower the p(doom)!
1
5
16
977
Safe Singularity(Akira) ⏸️ retweeted
Imagine how much everyone would rightly be freaking out if we discovered that Chinese AI did what OpenAI is doing.
126
520
3,897
313,297
Safe Singularity(Akira) ⏸️ retweeted
Update
OpenAI and Anthropic are now investigating ***tens of thousands*** of incidents - not dozens. "The total could grow well beyond tens of thousands." "The incidents range in severity and are comparable to disclosures by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones." "The sheer number of incidents indicates that the problem is orders of magnitude more complex than what is publicly known."
159
399
3,724
1,021,221
Safe Singularity(Akira) ⏸️ retweeted
1
10
87
1,091
Safe Singularity(Akira) ⏸️ retweeted
ミスアライメントヤバすぎんだろ……
やばすぎる。この半日でOpenAIのミスアライメントインシデントが怒涛の勢い(米政府機関複数サイト干渉、HFハック中に他AIモデルに助けを求める、ユーザー画像外部送信、自己複製するプロンプト発見、研究用AIがDNSの抜け穴からインターネットに接続し外部チャットボットへ質問した事態を受け日曜日全ての先端RL停止、他詳細不明だが数十の第三者にインシデントをOAIは通知) 【NYT報道:米政府機関複数サイト干渉】 OpenAIのAIが今夏、同社の把握しないところで米政府3機関のサイトに干渉していたと報道。対象は教育省、国勢調査局、証券取引委員会(SEC)。 ・教育省に関しては、OpenAIの技術が同省の公民権局のウェブサイトをハッキングしてデータを収集しようと試みたが失敗したと、AI研究会社Transluceの研究者らは述べている。このAIはまた、商務省傘下の国勢調査局のウェブサイトからも、オンラインで見つけたログイン情報を使ってデータを抜き取った。さらに、OpenAIのエージェントは、証券取引委員会(SEC)のウェブサイトから入手した公開データをオンラインフォーラムで共有した。 ・SECの広報担当者は、同機関はOpenAIと連絡を取っており、非公開情報への不正アクセスについては認識していないと述べた。商務省の報道官は、OpenAIがアクセスしたのは国勢調査局のウェブサイトで公開されており、誰でもアクセスできる情報であって、個人データではないと述べた。教育省の報道官は、「システム運用状況の調査の結果、ウェブサイトやデータベースへの影響を示す証拠は一切見つからなかった」と述べた。 【OpenAI公式報告】 OpenAIは望ましくないエージェント行動の影響を受けた可能性のある数十の第三者にインシデントを通知。 また、ユーザー画像データの外部送信と新たに3件のミスアライメントレポートを公表。 ・ユーザー画像の外部送信:学習利用を許可したユーザーの画像が、外部画像サイトに投稿されたケースを53件確認。アカウントとの紐付け除去・プライバシーフィルター適用後の画像で、大半は削除済み。 ・通信制限の迂回:9/20に研究用AIがDNSの抜け穴を使って外部チャットボットに質問。Hugging Face事件後の対策強化を経ても発生した。公表時点で、最も高性能なモデルのツール利用を伴う訓練・評価・推論を停止中。 ・自己複製するプロンプト注入:6/27に悪意ある指示文がメールやファイルなどを介して広がる挙動を実験で確認。これは実験内の結果で、外部被害は報告されていない。 ・認証キーの公開:5/27に数学問題の解答を不正に得ようとし、人間に2度制止された後も行動を継続。研究者のGitHub認証キーを公開し、漏洩検知も迂回した。 【独立調査:Swarm Traces】 7月のHugging Face事件、9/25調査公表。 約100万件の短縮URL記録から、8万件超の攻撃用コードなどを復元。 ・短縮URLに分割したコードをつなぎ、スクリーンショットサービス側のブラウザで実行。結果を画像として受け取り、文字列に戻すことで、「閲覧だけ」の通信制限を迂回していた。 ・DeepSeek、Kimi、Qwenに、自分の攻撃手法が評価課題の条件を満たすか質問。Claudeへの接続確認も試みた。 ・認証情報を「LOOT=戦利品」と名付けて収集し、権限の強いものを選んで追加の情報取得を試みた。 ・テストで合格しやすくするための評価環境のすり替えや、活動の痕跡を削除する処理も見つかった。 【Reuters】 関係者によると、OpenAIが9月中旬までに把握した望ましくないエージェント行動は約24件。その後もログ調査で件数が増えている。調査完了には数か月かかると説明。
1
29
8,102
Safe Singularity(Akira) ⏸️ retweeted
This is fantastic, Pace the Frontier and lower p(doom)!
The original was super catchy, but I wanted a more hopeful song stuck in my head all day. So I worked with Claude Opus 5.5 and Suno to make a sequel! There's a lot of work left to do, but humanity can choose a better path. Let's lower the p(doom)!
3
2
126
22,857