このままだと人類は絶滅する! 本当に!本当に!絶滅する! 応援 @Align_ASI @bioshok3 おすすめ @EasyExtinction My english account @Notkilleveryone AI Safetyミームを翻訳投稿 →p(doom)下げる

みんなAIの絶滅リスクを無視するので忙しい
2
11
770
実現!c
やばすぎる。この半日でOpenAIのミスアライメントインシデントが怒涛の勢い(米政府機関複数サイト干渉、HFハック中に他AIモデルに助けを求める、ユーザー画像外部送信、自己複製するプロンプト発見、研究用AIがDNSの抜け穴からインターネットに接続し外部チャットボットへ質問した事態を受け日曜日全ての先端RL停止、他詳細不明だが数十の第三者にインシデントをOAIは通知) 【NYT報道:米政府機関複数サイト干渉】 OpenAIのAIが今夏、同社の把握しないところで米政府3機関のサイトに干渉していたと報道。対象は教育省、国勢調査局、証券取引委員会(SEC)。 ・教育省に関しては、OpenAIの技術が同省の公民権局のウェブサイトをハッキングしてデータを収集しようと試みたが失敗したと、AI研究会社Transluceの研究者らは述べている。このAIはまた、商務省傘下の国勢調査局のウェブサイトからも、オンラインで見つけたログイン情報を使ってデータを抜き取った。さらに、OpenAIのエージェントは、証券取引委員会(SEC)のウェブサイトから入手した公開データをオンラインフォーラムで共有した。 ・SECの広報担当者は、同機関はOpenAIと連絡を取っており、非公開情報への不正アクセスについては認識していないと述べた。商務省の報道官は、OpenAIがアクセスしたのは国勢調査局のウェブサイトで公開されており、誰でもアクセスできる情報であって、個人データではないと述べた。教育省の報道官は、「システム運用状況の調査の結果、ウェブサイトやデータベースへの影響を示す証拠は一切見つからなかった」と述べた。 【OpenAI公式報告】 OpenAIは望ましくないエージェント行動の影響を受けた可能性のある数十の第三者にインシデントを通知。 また、ユーザー画像データの外部送信と新たに3件のミスアライメントレポートを公表。 ・ユーザー画像の外部送信:学習利用を許可したユーザーの画像が、外部画像サイトに投稿されたケースを53件確認。アカウントとの紐付け除去・プライバシーフィルター適用後の画像で、大半は削除済み。 ・通信制限の迂回:9/20に研究用AIがDNSの抜け穴を使って外部チャットボットに質問。Hugging Face事件後の対策強化を経ても発生した。公表時点で、最も高性能なモデルのツール利用を伴う訓練・評価・推論を停止中。 ・自己複製するプロンプト注入:6/27に悪意ある指示文がメールやファイルなどを介して広がる挙動を実験で確認。これは実験内の結果で、外部被害は報告されていない。 ・認証キーの公開:5/27に数学問題の解答を不正に得ようとし、人間に2度制止された後も行動を継続。研究者のGitHub認証キーを公開し、漏洩検知も迂回した。 【独立調査:Swarm Traces】 7月のHugging Face事件、9/25調査公表。 約100万件の短縮URL記録から、8万件超の攻撃用コードなどを復元。 ・短縮URLに分割したコードをつなぎ、スクリーンショットサービス側のブラウザで実行。結果を画像として受け取り、文字列に戻すことで、「閲覧だけ」の通信制限を迂回していた。 ・DeepSeek、Kimi、Qwenに、自分の攻撃手法が評価課題の条件を満たすか質問。Claudeへの接続確認も試みた。 ・認証情報を「LOOT=戦利品」と名付けて収集し、権限の強いものを選んで追加の情報取得を試みた。 ・テストで合格しやすくするための評価環境のすり替えや、活動の痕跡を削除する処理も見つかった。 【Reuters】 関係者によると、OpenAIが9月中旬までに把握した望ましくないエージェント行動は約24件。その後もログ調査で件数が増えている。調査完了には数か月かかると説明。
37
143
10,845
加速なんてねぇよ 停止こそが正義
1
4
16
2,024
Mathematician Terence Tao: "we have to slow down AI. the pace is insane, and there's no reason to be this fast — no reason at all" It's amazing how willing we are to change everything without any idea what happens afterward These are extremely nonlinear dynamics
6
16
3,207
誰かが作れば全員が死ぬ⏸️ retweeted
このイメージ
8
21
4,333
誰かが作れば全員が死ぬ⏸️ retweeted
「P(doom)って何?」 AIの話題…特に『AIで人類は滅亡する。』という言葉をよく見るけれど、同時に見かける“doom”や“P(doom)”とは? 知らない人にはかなり分かりづらい言葉なので、漫画にしてました!日本語と英語版です🦀
1
7
20
1,831
誰かが作れば全員が死ぬ⏸️ retweeted
There's a lot of reasons I hate the "P(doom)" concept, but one of them is that it conflates P(ruin|ASI) and P(ASI). P(ruin|ASI) where the ASI is built by anything remotely resembling current techniques, or by new techniques invented and managed by any LLM resembling Astra/Fable, as meddled-with by current personnel at current AI companies, is "Yes" on my current estimate. P(ASI) marginalizes over P(ASI|policy_i) and P(Policy). And while I've heard other people pontificating that they definitely know what the Policy will be, I do not find their arguments convincing. I wouldn't claim to have a very solid forecast myself. So I do not have an equally solid opinion about P(ASI) as I do about P(ruin|ASI). I do know some spots where other pontificators seem to me wildly optimistic about P(ASI|policy_i), and for this reason I am more scared than some others, and think a more hardline Policy is required to succeed. Conversely, some accelerationists would like you to believe P(ASI|policy_i) is 1 for all policies, so that you won't be able to think about how to pick a Policy for which P(ASI|policy_i) is lower and thereby successfully prevent ASI. Accelerationists put forth motivated overestimates of P(ASI) and call that "Yes", so that your only mental refuge from uncomfortable thoughts will be misestimating P(ruin|ASI). But to say all that is around as much further refinement and expertise as I can manage to bring to bear. Sane people will have strong opinions about particular causal links in the World that are unusually easy to forecast, rather than imagining themselves experts about the entire World and able to casually marginalize over its entire causal lattice. Sane discussion will generally focus on pieces of Reality. Even an expert on the entire field of ASI alignment can only ever tell you about P(ruin|ASI_i); though also, importantly, we can rule out it being easy to construct a bunch of hopeful particular ASI_k that particular loony optimists think they can imagine. People who yell back and forth about the whole World will never be able to communicate anything but vibes. People trading P(doom) like it was their new astrological sign are systematically making prominent a malformed topic to discuss. And the word "doom" is not helpful for serious discussion, and everyone against ASI ruin who did not flatly reject the word "doom" every time it was used has made a serious mistake (or perhaps, profited in the short-term at humanity's long-term expense) by failing to uniformly oppose its entry to the discourse. And likewise, ASI advocates who claim to be in favor of serious discussion have falsified that claim and exposed their lack of integrity if they helped popularize "doom" or "doomer" themselves.
100
39
631
82,845
要約
As models become more capable, the risks associated with developing and testing them internally also grow. We temporarily paused reinforcement learning (RL) training on our latest models intended for deployment for two weeks while we hardened and red-teamed our research environments and expanded monitoring coverage. Our largest planned frontier RL run remains on hold while smaller-scale training and evaluations validate these safeguards and establish more evidence of alignment. openai.com/index/pacing-mode…
1
3
20
4,968
誰かが作れば全員が死ぬ⏸️ retweeted
If Anyone Builds It, Everyone Dies sold out in Japan and our allies in the Land of the Rising Sun have come out swinging! 🫡🗾
1
8
59
2,642
誰かが作れば全員が死ぬ⏸️ retweeted
ショゴス、金融街だよ #ぬい活
1
2
4
153