Jinesis Lab led by Prof @ZhijingJin at @UofTCompSci @VectorInst conducts frontier research on Responsible AI, LLMs, and Causality.

Jinesis Lab (UToronto) retweeted
We are excited to share that we received over 420 submissions for the AI4GOOD workshop, and the review process has already started !! To keep things fair, we follow the strict rules of the main NeurIPS conference. We cannot accept any papers outside of the official OpenReview portal, and we cannot allow resubmissions or changes after a desk-rejection. Because we are managing so many papers, we won't be able to reply to individual queries.
We are so excited to host the 2nd edition of the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris, France πŸ‡«πŸ‡· We aim to bring together the AI safety, AI for social good, and AI policy/governance communities to connect what models can do as individual systems with what they do when deployed across populations. For full details, welcome to visit our website at: πŸ”— trustworthy-ai-for-good.gith… πŸ“December 12, 2026 Β· Paris πŸ‡«πŸ‡· We have an amazing lineup of speakers and plans. More details 🧡
2
16
2,954
πŸŽ‰Very excited to share that Jinesis Lab at the University of Toronto has 10 papers accepted at EMNLP 2026 β€” 5 at the main conference, 4 at Findings, and 1 at the Demo track! Our work this cycle spans interpretability, AI safety, computer vision, and multi-agent systems: πŸ“ŒMain Conference: β€’ Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders β€” arxiv.org/abs/2511.10840 β€’ PruneGround: Plug-and-play Spatial Pruning for 3D Visual Grounding β€” arxiv.org/abs/2606.31148 β€’ How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution β€’ Computation Graph Recovery from Chain-of-Thought β€” openreview.net/pdf?id=OHES5c… β€’ Preserving Historical Truth: Detecting Historical Revisionism in Large Language Models β€” arxiv.org/abs/2602.17433 πŸ“ŒFindings: β€’ Evaluating Second-Order Bias of LLMs Through Epistemic Entitlement β€” arxiv.org/abs/2606.17506 β€’ How Does Alignment Tuning Shape Representations of Sycophancy and Related Cue-Induced Biases in LLMs? β€” arxiv.org/pdf/2607.18114 β€’ Simulating Democratic Deliberation: Representing Pluralistic Preferences through Electoral Systems in Multi-Agent LLMs β€’ Fluid Representations in Reasoning Models β€” arxiv.org/abs/2602.04843 πŸ“ŒDemo Track: β€’ CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs Huge congratulations to all the authors and collaborators across the lab and our partner institutions. @AHarrasse1906, @flodraye, @psyonp, @ZhijingJin, @bschoelkopf, Duc Cao Dinh, @leduckhai, Chris Ngo, @TerryJCZhang, @vedantpalit1008, @RoderickWu4, Aydin Javadov, @francescortu, @JoeunYk05, @albecazzaniga, @radamihalcea, @RamaravindM, Raiyan Ahmed, Shion Guha, Syed Ishtiaque Ahmed, Prakhar Gupta, @_rfaulk, @SSzufa, Daniel Hoyer, Roland Bouffanais, @jzl86, Dmitrii Kharlapenko, @iarthsingh, @alesstolfo, @ArthurConmy, @mrinmayasachan, @TonyWu1105, @Jiarui_Liu_, Chih-Hao Hsu. #EMNLP2026 #NLP #Interpretability #AISafety #MachineLearning
2
1
25
5,674
Jinesis Lab (UToronto) retweeted
What is chain-of-thought actually doing inside the model? Our new #EMNLP2026 Findings paper shows reasoning models rebuild the meaning of words in activation space over a single trace, and that these representations causally drive performance. Thread belowπŸ‘‡
Our paper "Fluid Reasoning Representations" was accepted to EMNLP 2026 Findings. The question we asked: when a reasoning model spends 15k tokens thinking, what is actually changing inside it? Short answer: it is rebuilding the meaning of words in activation space, mid-trace.
1
4
75
5,665
Jinesis Lab (UToronto) retweeted
🚨 8 days left to submit to the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris We’re still accepting submissions. Join us in bridging AI safety, social good, and real-world impact, so that more capable models actually help society at scale. Our workshop is non-archival, and we will recognize outstanding papers and top reviewers with awards. Deadline: August 29, 2026 (AoE) More details and submission β†’ trustworthy-ai-for-good.gith… πŸ“§ Sponsorship & questions: zjingchen@cs.toronto.edu Let's bridge trustworthy AI and real-world impact. We hope to see you all in Paris. @TerryJCZhang @ChanglingXavier @_agirlyengineer @ozzaney0101, He (Shawn) Shuang, Jerick Shi, Prakhar Gupta, Kexin Li (Cassie), Wenjun (Wendy) Qiu @ZhijingJin @radamihalcea @MilindTambe_AI @david_lie @casdewitt @VectorInst @JinesisLab @EuroSafeAI @MPI_IS @UofTCompSci @TorontoSRI @CIFAR_News @ELLISInst_Tue @UMichCSE @michigan_AI @Harvard @ETH_en @CarnegieMellon @UniofOxford @NeurIPSConf @UofT
We are so excited to host the 2nd edition of the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris, France πŸ‡«πŸ‡· We aim to bring together the AI safety, AI for social good, and AI policy/governance communities to connect what models can do as individual systems with what they do when deployed across populations. For full details, welcome to visit our website at: πŸ”— trustworthy-ai-for-good.gith… πŸ“December 12, 2026 Β· Paris πŸ‡«πŸ‡· We have an amazing lineup of speakers and plans. More details 🧡
1
5
16
4,026
Jinesis Lab (UToronto) retweeted
We are so excited to host the 2nd edition of the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris, France πŸ‡«πŸ‡· We aim to bring together the AI safety, AI for social good, and AI policy/governance communities to connect what models can do as individual systems with what they do when deployed across populations. For full details, welcome to visit our website at: πŸ”— trustworthy-ai-for-good.gith… πŸ“December 12, 2026 Β· Paris πŸ‡«πŸ‡· We have an amazing lineup of speakers and plans. More details 🧡
2
15
63
19,904
Jinesis Lab (UToronto) retweeted
πŸ“‘Announcing NewInML workshop at #NeurIPS2026 We are so excited to host the 8th edition of the NewInML (@NewInML) workshop, which will take place in Paris, France πŸ‡«πŸ‡· We are here to help you out gain confidence and network with the community! If you are new to ML research or have limited publishing experience, it's a perfect opportunity for you to submit your first work, get high quality reviews from reviewers and also meet the experts, and of course, enjoy beautiful Paris πŸ₯ We are delighted to have an amazing speaker line-up and interesting plans. More details 🧡
2
15
60
19,532
Jinesis Lab (UToronto) retweeted
We're delighted to host the Trustworthy AI for Good (AI4GOOD) workshop at @NeurIPSConf in Paris πŸ‡«πŸ‡· πŸ“’ We're recruiting reviewers, plus an award for top reviewers πŸ† If you're working on AI Safety, Alignment, AI for Science or related topics, we're looking for you. Feel free to express your interest here: forms.gle/HpkPoE8HqMHH1geL9
🌟 We’d love to welcome you as a reviewer for the 2nd Trustworthy AI4GOOD Workshop at #NeurIPS2026 in Paris! πŸ‡«πŸ‡· Following the success of our first edition at #ICML2026, this workshop will once again bring together researchers and practitioners working to ensure that advances in AI are not only powerful, but also trustworthy, responsible, and genuinely beneficial to society. Each reviewer will be assigned no more than 3 papers through OpenReview, with no paper bidding or rebuttal phase. πŸ—“οΈ Review period: September 2026 ⏰ Review deadline: September 22, 2026, AOE Please make sure your OpenReview profile is up to date before registering. Join us as a reviewer: forms.gle/HpkPoE8HqMHH1geL9 We’d be delighted to have you help shape this year’s workshop! πŸŽ‰ #AI4GOOD #TrustworthyAI #NeurIPS #AISafety
4
19
2,839
🌟 We’d love to welcome you as a reviewer for the 2nd Trustworthy AI4GOOD Workshop at #NeurIPS2026 in Paris! πŸ‡«πŸ‡· Following the success of our first edition at #ICML2026, this workshop will once again bring together researchers and practitioners working to ensure that advances in AI are not only powerful, but also trustworthy, responsible, and genuinely beneficial to society. Each reviewer will be assigned no more than 3 papers through OpenReview, with no paper bidding or rebuttal phase. πŸ—“οΈ Review period: September 2026 ⏰ Review deadline: September 22, 2026, AOE Please make sure your OpenReview profile is up to date before registering. Join us as a reviewer: forms.gle/HpkPoE8HqMHH1geL9 We’d be delighted to have you help shape this year’s workshop! πŸŽ‰ #AI4GOOD #TrustworthyAI #NeurIPS #AISafety
1
7
32
12,905
πŸŽ‰ With the successful conclusion of our 1st Trustworthy AI4GOOD workshop at #ICML2026, we are delighted to share that its 2nd edition has been accepted for #NeurIPS2026 in Paris! πŸŽ‰πŸ‡«πŸ‡· The workshop selection process was highly competitive: only 22% of 454 workshop proposals were accepted. We are grateful for the opportunity to bring the community together again! πŸ™Œ Paris, here we come! πŸ‡«πŸ‡· ICML speakers: @Yoshua_Bengio, @OanaIgnatRo, @jzl86, @maksym_andr, Jenny Ni and @NamGoPro. Panel: @MilindTambe_AI, @lrhammond, Gopal Sarma and @ARGleave. Organizers: @TerryJCZhang, @ZhijingJin, @radamihalcea, @MilindTambe_AI, David Lie, @_agirlyengineer, @davidguzman1120, @ChanglingXavier, @Jerick1380, Prakhar Gupta, @vantru0ng, @EttoreGran, He (Shawn) Shuang, @ozzaney0101, Kexin Li and Wenjun Qiu. Thank you for making our first edition unforgettable!πŸ™ Thank you for making AI4GOOD possible! πŸ’™ trustworthy-ai-for-good.gith…
1
8
34
5,914
Our Jinesis Lab at the University of Toronto is bringing three exciting papers to #COLM2026! πŸŽ‰ These projects push the frontiers of #AISafety, #AgentSecurity, and #AIForScience.πŸ€– Huge congratulations to all collaborators and co-authors! Papers: πŸ›‘οΈODILE: Orthogonal Disruption of Injected Tool-Call Embeddings for Agentic Prompt Injection Defense πŸ”­Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints ⚠️One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety @UofT,@VectorInst,@EuroSafeAI,@ETH_en,@MPI_IS,@UMich
5
20
1,317
🌟 What an incredible experience at #ICML2026! 🌟 Our team at Jinesis Lab had a fantastic time presenting research across several exciting and important areas: 🀝 Multi-agent LLMs 🧠 Causal reasoning πŸ›‘οΈ AI safety We were proud to share this work alongside our collaborators and to engage with so many researchers from across the machine learning community. πŸ’‘ Just as valuable as presenting our own work was learning from the research others brought to ICML. From thought-provoking talks and innovative poster presentations to insightful discussions and spontaneous hallway conversations, the conference was full of ideas that challenged our thinking and opened new directions for future work. ✨ We are grateful to everyone who attended our presentations, visited our posters, shared feedback, and took the time to discuss their own research with us. πŸ™Œ A huge thank-you to our lab members, collaborators, and the broader ICML community for making the experience so memorable. We are leaving inspired, energized, and excited to build on the ideas and connections formed throughout the conference. πŸŒπŸ”¬πŸš€ #ICML #MachineLearning #ArtificialIntelligence #MultiAgentSystems #LLM #CausalReasoning #AISafety #ResponsibleAI #AIResearch #ResponsibleAI #AIResearch
5
29
2,632
πŸŽ‰βœ¨ The Trustworthy AI for Good Workshop at ICML 2026 was a huge success! βœ¨πŸŽ‰ We are incredibly grateful to everyone who joined us in Seoul πŸ‡°πŸ‡· for a full day of inspiring research, thoughtful discussions, and meaningful connections around building AI systems that are both trustworthy and socially beneficial. πŸ€–πŸŒπŸ’‘ A heartfelt thank-you to our incredible keynote speakers: 🎀 @Yoshua_Bengio, @OanaIgnatRo, @jzl86, @maksym_andr, Jenny Ni, Naman Goyal Your insights, perspectives, and thought-provoking ideas made the workshop truly memorable. 🌟🧠 A special thank-you to our outstanding panelists: πŸ’¬ @MilindTambe_AI, @lrhammond, Gopal Sarma, @ARGleave And to our wonderful moderator, @ZhijingJin, for leading such an engaging and insightful conversation. πŸŽ™οΈ We are also deeply grateful to all our oral and poster presenters, reviewers, organizers, volunteers, sponsors, and every attendee who contributed questions, ideas, feedback, and enthusiasm throughout the day. πŸš€ The energy of this community reinforced the importance of bringing together: πŸ›‘οΈ AI safety 🌍 AI for social good βš–οΈ AI policy and governance 🀝 Cooperative AI πŸ“° Information integrity πŸ›οΈ Civic discourse Thank you for making Trustworthy AI for Good @ #ICML2026 such a memorable and impactful event! 🎊 We look forward to continuing the conversation, strengthening this community, and building on the connections formed at the workshop. 🀝 #ICML2026 #AIForGood #TrustworthyAI #AISafety #ResponsibleAI #MachineLearning #CooperativeAI #AIResearch #TechForGood #ArtificialIntelligence
1
3
24
1,936
🚨 Thrilled to share that our lab will be presenting the πŸ† Best Paper at the NExT-Game Workshop at #ICML2026 today! 🎀 When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games πŸ† Best Paper @ NExT-Game Workshop πŸ“ Conference Room S307 πŸ“… Fri, Jul 10 πŸ• 13:00–13:20 KST Authors: @JerickShi @TerryJCZhang @bschoelkopf @conitzer @ZhijingJin πŸ€– We introduce a three-stage endogenous promise protocol for repeated multi-agent games that asks not only whether LLM agents honor their public commitments when they can privately deviate, but also how model-on-model composition influences premeditated deception and persistent exploitation. πŸ“Š Across six canonical games spanning binary and numerical action spaces, our evaluation of frontier models (GPT-5.2, Llama-4-Maverick, Claude-Opus-4.6) reveals: πŸ”Ή Over 90% of promise-breaking instances are premeditated in agents' private plans. πŸ”Ή Mixed-model groups with mismatched communication frameworks create systemic, persistent payoff gaps of up to 5.00 points from Round 0. πŸ“„ Paper: openreview.net/forum?id=v8nY… #MultiAgentSystems #LLMs #GameTheory #AI #ICML2026
1
2
8
1,139
Excited to share more of our workshop papers today, July 11, at #ICML2026!πŸ“πŸ‡°πŸ‡·πŸ“· πŸ”§AIWILD (⭐Oral⭐) AF-ARENA: A Multi-Dimensional Evaluation Suite for Alignment Faking πŸ—“οΈ 10:10–10:40 AM KST πŸ“Hall B2. πŸ”—openreview.net/forum?id=vFqn… πŸ”§AI4Math Proving Your Way to Cooperation: Formalizing Proof-Based Open Source Game Theory in Lean πŸ“Room HALL D1 πŸ”—openreview.net/forum?id=Wc5T… πŸ”§AI4Physics Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints πŸ“Conference Room S402 πŸ”—aips-uoft.github.io/Stargaze… πŸ”§NexT-Game GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory πŸ“Conf. room S307 πŸ”—openreview.net/forum?id=F0Dz… πŸ”§Pluralistic Alignment The Wedge Questions: Latent Cultural Boundaries in LLMs via Persona Projection Divergence πŸ“Room 403, Board #4613 πŸ”—openreview.net/forum?id=zHQu… Huge thanks to all collaborators, co-authors, organizers, and workshop communities β€” looking forward to many great discussions at #ICML πŸš€
4
243
Thrilled to share that our lab will present the πŸ†Best PaperπŸ†of RLxF Workshop @ #ICML2026 today! πŸ“„ Transferability for General Reasoning: An Automated Curriculum for Multi-Domain LLM RL πŸ“ GRAND BALLROOM 101-102 πŸ—“οΈ Fri, Jul 10 KST πŸ•š 4:00 PM – 5:00 PM KST πŸŽ“ Authors: @_yongjinny, @Jiarui_Liu_, @yinghui_he_, @leczhang, @bschoelkopf, and @ZhijingJin We introduce Transfer-Aware Curriculum (TAC), an automated curriculum for multi-domain RLVR that asks not only which domain is learnable right now, but also which domain’s gradients transfer well to the rest of the reasoning suite. 🧩 Across math, code, logic, simulation, table reasoning, and STEM, TAC achieves the best macro-averaged accuracy on both Qwen3-1.7B and Llama3.2-3B, improving over strong curriculum baselines with <1% wall-clock overhead. πŸ“ Huge thanks to all collaborators, reviewers, and the RLxF organizers β€” and please stop by to chat if you’re at #ICML2026! πŸš€ πŸ“„ Paper: arxiv.org/pdf/2606.25178 πŸ“Ž Code: github.com/YangYongJin/trans… #ICML #ICML2026 #RLxF #ReinforcementLearning #LLM #Reasoning #RLVR
4
15
1,111
Excited to share our workshop papers being presented TODAY at #ICML2026 in Seoul πŸ‡°πŸ‡·βœ¨ Please stop by the posters and chat with us! πŸ”§ FoGen πŸ“„ Test of Time: Rethinking Temporal Signal of Benchmark Contamination πŸ—“οΈ 9:10–10:00 AM, 2:00–2:50 PM KST πŸ“ ROOM 318 🀝 AI4GOOD 🌟Spotlight🌟 πŸ“„ Multi-Agent AI Systems Need Institutional Design, Not Just Model-Level Alignment πŸ—“οΈ 11:25 AM – 1:00 PM KST πŸ“ GRAND BALLROOM 103 πŸ”¬ AI4Research πŸ“„ Causal AI Scientist: Towards End-to-End Causal Inference with Large Language Models πŸ—“οΈ 11:50 AM – 1:00 PM KST πŸ“ AUDITORIUM 🚨 FAGEN πŸ“„ What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs πŸ—“οΈ 12:00–2:00 PM KST πŸ“ GRAND BALLROOM 104-105 🀝 AI4GOOD πŸ“„ Mechanism Design Is Not Enough: Prosocial Agents for Cooperative πŸ“„ CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas πŸ“„ Weight-Level Defenses Improve LLM Agent Adversarial Robustness πŸ“„ Evaluating Cooperation in LLM Social Groups through Elected Leadership πŸ“„ Multi-Agent AI Systems Need Institutional Design, Not Just Model-Level Alignment πŸ—“οΈ 4:20–5:00 PM KST πŸ“ GRAND BALLROOM 103 Huge thanks to all collaborators, co-authors, organizers, and workshop communities β€” looking forward to many great discussions at #ICML2026 πŸš€
2
7
536
Interested in AI safety, AI for social good, and AI policy/governance? Join us at the #ICML2026 Workshop: Trustworthy AI for Good! πŸ“… Friday, 10 July 2026, 08:00-17:00 πŸ“· Grand Ballroom 103 (near registration area) See you there! πŸ‘ Congratulations to first authors from the oral papers: @yaowenye123, @EmanuelTewolde, Hanbo Huang, Jennifer Za, @narutatsuri, @CuhelJan, Xu Lui
3
17
8,101
πŸ“£ We are presenting 6 main conference papers πŸš€and 14 workshop papers (including πŸ†2 Best PapersπŸ†) at #ICML2026 in Korea! Also hosting one of the largest workshops, Trustworthy AI for Good, on July 10th 🌍❀️. We push the frontiers on #AISafety, #MultiAgent, and #CausalReasoning at @JinesisLab! πŸŽ‰ Huge congratulations to all collaborators and co-authors. Excited to discuss these projects in Seoul! Feel free to reach out and talk to our 20+ members and collaborators in Korea @ZhijingJin, @_AndreiMuresanu, @iarthsingh, @ChanglingXavier, @davidguzman1120, @EmanuelTewolde, @ettogran, @FurkanDanismann, @Jerick1380, @PepijnCobben, @rishit_dagli, @_rfaulk, @SimkoSamuel, @TerryJCZhang, @vantru0ng, @zhxiao03, @x_angelohuang, @yahang_qi, @ozzaney0101, @_yongjinny. Happy for collaboration on any of the above topics 🀝 EuroSafeAI, University of Toronto, ETH ZΓΌrich, Max Planck Institute for Intelligent Systems Main conference spotlight 🌟 Safe Models Do Not Guarantee Safe Societies: The Case for Sociopolitical Risk Main conference posters πŸ“Œ CauSciBench: Can LLMs Automate Causal Inference in Real-World Scientific Research? πŸ“Œ Training with Honeypots: Reshaping How LLMs Fail πŸ“Œ CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas πŸ“Œ Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution πŸ“Œ LLM for Physics Research Requires Domain-Specialized Training and Tooling Workshop best papers πŸ†When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games BEST PAPER@NExT-Game Workshop πŸ†Transferability for General Reasoning: An Automated Curriculum for Multi-Domain LLM RL BEST PAPER@RLxF Workshop Workshop oral and spotlight 🎀 AF-ARENA: A Multi-Dimensional Evaluation Suite for Alignment Faking β€” AIWILD 🌟 Multi-Agent AI Systems Need Institutional Design, Not Just Model-Level Alignment β€” AI4GOOD Workshop papers πŸ“„ The Wedge Questions: Latent Cultural Boundaries in LLMs via Persona Projection Divergence β€” Pluralistic Alignment πŸ“„ GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theory β€” NExT-Game πŸ“„ Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints β€” AI4Physics πŸ“„ Test of Time: Rethinking Temporal Signal of Benchmark Contamination β€” FoGen πŸ“„ CoopEval: Benchmarking Cooperation-Sustaining Mechanisms and LLM Agents in Social Dilemmas β€” AI4GOOD πŸ“„ Weight-Level Defenses Improve LLM Agent Adversarial Robustness β€” AI4GOOD πŸ“„ Evaluating Cooperation in LLM Social Groups through Elected Leadership β€” AI4GOOD πŸ“„ Causal AI Scientist: Towards End-to-End Causal Inference with Large Language Models β€” AI4Research πŸ“„What Game-Theoretic Benchmarks Miss: Strategic Silence in Multi-Agent LLMs β€” FAGEN πŸ“„Proving Your Way to Cooperation: Formalizing Proof-Based Open Source Game Theory in Lean β€” AI4Math
2
16
5,428