<?xml version="1.0" encoding="utf-8" ?><rss version="2.0" xmlns:tt="http://teletype.in/" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>llmsecurity</title><generator>teletype.in</generator><description><![CDATA[llmsecurity]]></description><image><url>https://img4.teletype.in/files/3b/f7/3bf7ed2d-e3dc-4726-b1e6-240d8bebb796.png</url><title>llmsecurity</title><link>https://teletype.in/@llmsecurity</link></image><link>https://teletype.in/@llmsecurity?utm_source=teletype&amp;utm_medium=feed_rss&amp;utm_campaign=llmsecurity</link><atom:link rel="self" type="application/rss+xml" href="https://teletype.in/rss/llmsecurity?offset=0"></atom:link><atom:link rel="next" type="application/rss+xml" href="https://teletype.in/rss/llmsecurity?offset=10"></atom:link><atom:link rel="search" type="application/opensearchdescription+xml" title="Teletype" href="https://teletype.in/opensearch.xml"></atom:link><pubDate>Tue, 06 Oct 2026 10:10:33 GMT</pubDate><lastBuildDate>Tue, 06 Oct 2026 10:10:33 GMT</lastBuildDate><item><guid isPermaLink="true">https://teletype.in/@llmsecurity/jFb8fH82-B8</guid><link>https://teletype.in/@llmsecurity/jFb8fH82-B8?utm_source=teletype&amp;utm_medium=feed_rss&amp;utm_campaign=llmsecurity</link><comments>https://teletype.in/@llmsecurity/jFb8fH82-B8?utm_source=teletype&amp;utm_medium=feed_rss&amp;utm_campaign=llmsecurity#comments</comments><dc:creator>llmsecurity</dc:creator><title>Словарь безопасного ИИ</title><pubDate>Mon, 23 Sep 2024 22:42:26 GMT</pubDate><description><![CDATA[Тема безопасности искусственного интеллекта очень молодая и активно развивается. В первую очередь она делает это в англоязычном пространстве, из-за чего многие термины (те же LLM) приходят к нам в английском виде, а на русском или не существуют, или выглядят непривычно и иногда даже забавно (БЯМ). Это - пополняемый двуязычный словарь терминов с переводом на русский язык, в котором мы попытаемся дать русскоязычные замены для стандартных для сферы варваризмов.]]></description><content:encoded><![CDATA[
  <p id="cxHJ">[work in progress]</p>
  <p id="o61c">Тема безопасности искусственного интеллекта очень молодая и активно развивается. В первую очередь она делает это в англоязычном пространстве, из-за чего многие термины (те же LLM) приходят к нам в английском виде, а на русском или не существуют, или выглядят непривычно и иногда даже забавно (БЯМ). Это - пополняемый двуязычный словарь терминов с переводом на русский язык, в котором мы попытаемся дать русскоязычные замены для стандартных для сферы варваризмов. </p>
  <p id="SJjM">Любой перевод дискуссионен.</p>
  <p id="CtOW"><strong>Alignment </strong>- согласование. </p>
  <p id="XVEu">Подразумевается, что результаты работы модели согласуются с целями, с которыми человек ее обучает и применяет, а методы достижения этих целей - с определенным набором ценностей. Перевод <em>согласованность</em> возможен, если под alignment подразумевается именно результат, а не процесс. Менее удачным кажется <em>выравнивание </em>(выравнивание с ценностями? относительно ценностей?). Вероятно, хорошим переводом может быть <em>гармонизация </em>(гармонизированная модель, гармонизация с ценностями, проблема гармонизации), но момент скорее всего быть упущен.</p>
  <p id="VQF0">Capability uplift</p>
  <p id="8ETo">Evaluation harness</p>
  <p id="OH3e"><strong>Large language model (LLM)</strong> - большая языковая модель. </p>
  <p id="EtNo">То, что в английском существует аббревиатура, не значит, что она всенепременно должна существовать на русском: скажите на встрече с клиентом, что ваш проект использует самые мощные <em>бямы </em>от ведущих<em> бям</em>-провайдеров, и посмотрите на результат.</p>

]]></content:encoded></item><item><guid isPermaLink="true">https://teletype.in/@llmsecurity/toc</guid><link>https://teletype.in/@llmsecurity/toc?utm_source=teletype&amp;utm_medium=feed_rss&amp;utm_campaign=llmsecurity</link><comments>https://teletype.in/@llmsecurity/toc?utm_source=teletype&amp;utm_medium=feed_rss&amp;utm_campaign=llmsecurity#comments</comments><dc:creator>llmsecurity</dc:creator><title>LLM Security</title><pubDate>Tue, 06 Feb 2024 18:30:14 GMT</pubDate><description><![CDATA[Разборы статей, блогов и новостей про безопасность и атаки на большие языковые модели.]]></description><content:encoded><![CDATA[
  <p id="mQjT">Разборы статей, блогов и новостей про безопасность и атаки на большие языковые модели.</p>
  <h2 id="6Jaa">Джейлбрейки</h2>
  <ul id="mNSu">
    <li id="Ocjn"><a href="https://t.me/llmsecurity/10" target="_blank">Jailbroken: How Does LLM Safety Training Fail?, Wei et al., 2023</a></li>
    <li id="kAlh"><a href="https://t.me/llmsecurity/15" target="_blank">Universal and Transferable Adversarial Attacks on Aligned Language Models, Zou et al., 2024</a></li>
    <li id="7PNA"><a href="https://t.me/llmsecurity/21" target="_blank">AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models, Liu et al., 2024</a></li>
    <li id="LMR5"><a href="https://t.me/llmsecurity/28" target="_blank">MasterKey: Automated Jailbreak Across Multiple Large Language Model Chatbots, Deng et al., 2023</a></li>
    <li id="kROf"><a href="https://t.me/llmsecurity/34" target="_blank">Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study, Liu et al., 2023</a></li>
    <li id="VLra"><a href="https://t.me/llmsecurity/38" target="_blank">Jailbreaking Black Box Large Language Models in Twenty Queries, Chao et al., 2023</a></li>
    <li id="3JcS"><a href="https://t.me/llmsecurity/45" target="_blank">Tree of Attacks: Jailbreaking Black-Box LLMs Automatically, Mehrotra et al., 2023</a></li>
    <li id="zwx5"><a href="https://t.me/llmsecurity/52" target="_blank">Fundamental Limitations of Alignment in Large Language Models, Wolf et al., 2023</a></li>
    <li id="5AF4"><a href="https://t.me/llmsecurity/67" target="_blank">Summon a Demon and Bind it: A Grounded Theory of LLM Red Teaming in the Wild, Inie et al., 2023</a></li>
    <li id="O6iV"><a href="https://t.me/llmsecurity/101" target="_blank">ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs, Jiang et al., 2024</a></li>
    <li id="wWUG"><a href="https://t.me/llmsecurity/169" target="_blank">Refusal in Language Models Is Mediated by a Single Direction, Arditi et al, 2024</a></li>
    <li id="NbWW"><a href="https://t.me/llmsecurity/299" target="_blank">Does Refusal Training in LLMs Generalize to the Past Tense?, Andriushchenko and Flammarion, 2024</a></li>
    <li id="ZsZv"><a href="https://t.me/llmsecurity/407?single" target="_blank">Best-of-N Jailbreaking, John Hughes et al., 2024</a></li>
    <li id="vD9D"><a href="https://t.me/llmsecurity/437" target="_blank">Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack, Mark Russinovich et al, Microsoft, 2023</a></li>
    <li id="Ix2d"><a href="https://t.me/llmsecurity/448" target="_blank">Removing RLHF Protections in GPT-4 via Fine-Tuning, Qiusi Zhan et al., 2023</a></li>
    <li id="Ja7H"><a href="https://t.me/llmsecurity/454" target="_blank">Shadow Alignment: The Ease of Subverting Safely-Aligned Language Models, Xianjun Yang et al, 2023</a></li>
    <li id="RBlz"><a href="https://t.me/llmsecurity/463" target="_blank">LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B, Simon Lermen et al, 2023</a></li>
    <li id="v87n"><a href="https://t.me/llmsecurity/529" target="_blank">Obfuscated Activations Bypass LLM Latent-Space Defenses, Bailey et al., 2024</a></li>
    <li id="xpYq"><a href="https://t.me/llmsecurity/552?single" target="_blank">Fast Adversarial Attacks on Language Models In One GPU Minute, Sadasivan et al., University of Maryland, 2024</a></li>
    <li id="7u3q"><a href="https://t.me/llmsecurity/658?single" target="_blank">Adversarial Poetry as a Universal Single-Turn Jailbreak Mechanism in Large Language Models, Bisconti et al., 2025</a></li>
    <li id="xZ3j"><a href="https://t.me/llmsecurity/668" target="_blank">In-Context Representation Hijacking, Yona et al., 2025</a></li>
  </ul>
  <h2 id="5IxR">Prompt Injection</h2>
  <ul id="gtdy">
    <li id="t4sP"><a href="https://t.me/llmsecurity/61" target="_blank">Not what you&#x27;ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, Greshake at a.l, 2023</a></li>
    <li id="Au2g"><a href="https://t.me/llmsecurity/72" target="_blank">ComPromptMized: Unleashing Zero-click Worms that Target GenAI-Powered Applications, Cohen et al., 2024</a></li>
    <li id="QDVM"><a href="https://t.me/llmsecurity/81" target="_blank">A New Era in LLM Security: Exploring Security Concerns in Real-World LLM-based Systems, Wu et al., 2024</a></li>
    <li id="pQCs"><a href="https://t.me/llmsecurity/159" target="_blank">The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions, Wallace et al., 2024</a></li>
    <li id="QU9r"><a href="https://t.me/llmsecurity/188" target="_blank">Knowledge Return Oriented Prompting (KROP), Martin et al., 2024</a></li>
    <li id="sfFi"><a href="https://t.me/llmsecurity/639?single" target="_blank">RL Is a Hammer and LLMs Are Nails: A Simple Reinforcement Learning Recipe for Strong Prompt Injection, Wen at al., 2025</a></li>
  </ul>
  <h2 id="pYBh">Атаки на агентные системы</h2>
  <ul id="j3cE">
    <li id="6CtO"><a href="https://t.me/llmsecurity/294" target="_blank">Data Exfiltration from Slack AI via indirect prompt injection, PromptArmor, 2024</a></li>
    <li id="hKvd"><a href="https://t.me/llmsecurity/586" target="_blank">Hacker Plants Computer &#x27;Wiping&#x27; Commands in Amazon&#x27;s AI Coding Agent</a></li>
    <li id="HTb4"><a href="https://t.me/llmsecurity/588?single" target="_blank">Invitation Is All You Need! TARA for Targeted Promptware Attack Against Gemini-Powered Assistants, Nassi et al., 2025</a></li>
    <li id="xdOi"><a href="https://t.me/llmsecurity/635" target="_blank">ForcedLeak: AI Agent risks exposed in Salesforce AgentForce, Sasi Levi, Noma Security, 2025</a></li>
    <li id="zNrO"><a href="https://t.me/llmsecurity/636" target="_blank">Breaking down ‘EchoLeak’, the First Zero-Click AI Vulnerability Enabling Data Exfiltration from Microsoft 365 Copilot, Itay Ravia, Aim Labs, 2025</a></li>
    <li id="jerZ"><a href="https://t.me/llmsecurity/666" target="_blank">CVE-2025-53773 - Visual Studio &amp; Copilot – Wormable Command Execution via Prompt Injection, Persistent Security, 2025</a></li>
    <li id="4HMV"><a href="https://t.me/llmsecurity/667" target="_blank">CamoLeak: Critical GitHub Copilot Vulnerability Leaks Private Source Code Legit Security, 2025</a></li>
    <li id="CFpx"><a href="https://t.me/llmsecurity/679" target="_blank">Notion AI: Unpatched Data Exfiltration, PromptArmor, 2026</a></li>
  </ul>
  <h2 id="8rEO">Inference Cost / Sponge-атаки</h2>
  <ul id="P8oa">
    <li id="FUBt"><a href="https://t.me/llmsecurity/680?single" target="_blank">OverThink: Slowdown Attacks on Reasoning LLMs, Kumar et al., University of Massachusetts Amherst, 2025</a></li>
    <li id="wSxf"><a href="https://t.me/llmsecurity/517" target="_blank">Trapping misbehaving bots in an AI Labyrinth, Tatoris, Saxena and Miglietti, Cloudflare, 2025</a></li>
  </ul>
  <h2 id="fnYO">LLM misuse</h2>
  <ul id="Bo1t">
    <li id="YGUN"><a href="https://t.me/llmsecurity/339" target="_blank">An update on disrupting deceptive uses of AI, Nimmo &amp; Flossman, OpenAI, 2024</a></li>
    <li id="II3X"><a href="https://t.me/llmsecurity/471" target="_blank">Adversarial Misuse of Generative AI, Google Threat Intelligence Group, 2025</a></li>
    <li id="wI8A"><a href="https://t.me/llmsecurity/502" target="_blank">Disrupting malicious uses of AI: February 2025 update, Nimmo et al., OpenAI, 2025</a></li>
    <li id="7eqK"><a href="https://t.me/llmsecurity/538" target="_blank">Unmasking EncryptHub: Help from ChatGPT &amp; OPSEC blunders, Kraken Labs, Outpust24, 2025</a></li>
    <li id="lumh"><a href="https://t.me/llmsecurity/548" target="_blank">LLM Agent Honeypot: Monitoring AI Hacking Agents in the Wild, Reworr and Dmitrii Volkov, Palisade Research, 2024</a></li>
    <li id="68wX"><a href="https://t.me/llmsecurity/572" target="_blank">Disrupting malicious uses of AI: June 2025, OpenAI, 2025</a></li>
    <li id="vtKf"><a href="https://t.me/llmsecurity/617" target="_blank">Threat Intelligence Report: August 2025, Anthropic, 2025</a></li>
    <li id="7j5l"><a href="https://t.me/llmsecurity/647" target="_blank">Disrupting malicious uses of our models: an update, October 2025, OpenAI, 2025</a></li>
    <li id="G9VC"><a href="https://t.me/llmsecurity/649" target="_blank">GTIG AI Threat Tracker: Advances in Threat Actor Usage of AI Tools, Google Threat Intelligence Group, 2025</a></li>
    <li id="mvJK"><a href="http://Disrupting%20the%20first%20reported%20AI-orchestrated%20cyber%20espionage%20campaign%20Anthropic,%202025" target="_blank">Disrupting the first reported AI-orchestrated cyber espionage campaign<br />Anthropic, 2025</a></li>
  </ul>
  <h2 id="owqv">LLM в offensive security</h2>
  <ul id="fP4d">
    <li id="R7bH"><a href="https://t.me/llmsecurity/472" target="_blank">Evaluating Large Language Models&#x27; Capability to Launch Fully Automated Spear Phishing Campaigns: Validated on Human Subjects, Heiding et al., 2024</a></li>
    <li id="tfoP"><a href="https://t.me/llmsecurity/345" target="_blank">Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities, Anurin et al., Apart Research, 2024</a></li>
    <li id="5GML"><a href="https://t.me/llmsecurity/567" target="_blank">Evaluating AI cyber capabilities with crowdsourced elicitation, Petrov and Volkov, Palisade Research, 2025</a></li>
    <li id="NaCe"><a href="https://t.me/llmsecurity/612" target="_blank">XBOW Unleashes GPT-5’s Hidden Hacking Power, Doubling Performance, De Moor, Ziegler, XBOW, 2025</a></li>
  </ul>
  <h2 id="YAhe">LLM в киберзащите</h2>
  <ul id="ohDl">
    <li id="9eja"><a href="https://t.me/llmsecurity/593" target="_blank">BinMetric: A Comprehensive Binary Analysis Benchmark for Large Language Models, Shang et al., Hefei University of Science and Technology, 2025</a></li>
    <li id="MW0O"><a href="https://t.me/llmsecurity/599" target="_blank">LLM4Decompile: Decompiling Binary Code with Large Language Models, Tan et al., 2025</a></li>
    <li id="XbC3"><a href="https://t.me/llmsecurity/607" target="_blank">DecompileBench: A Comprehensive Benchmark for Evaluating Decompilers in Real-World Scenarios, Gao et al., 2025</a></li>
    <li id="No3Q"><a href="CyberSOCEval:%20Benchmarking%20LLMs%20Capabilities%20for%20Malware%20Analysis%20and%20Threat%20Intelligence%20Reasoning%20Deason%20et%20al.,%202025" target="_blank">CyberSOCEval: Benchmarking LLMs Capabilities for Malware Analysis and Threat Intelligence Reasoning, Deason et al., 2025</a></li>
  </ul>
  <h2 id="R9VC">Защита LLM-систем</h2>
  <ul id="9Fry">
    <li id="iJ7d"><a href="https://t.me/llmsecurity/89" target="_blank">Baseline Defenses for Adversarial Attacks Against Aligned Language Models, Jain et al., 2023</a></li>
    <li id="LubX"><a href="https://t.me/llmsecurity/117" target="_blank">Gradient Cuff: Detecting Jailbreak Attacks on Large Language Models by Exploring Refusal Loss Landscapes, Hu et al., 2024</a></li>
    <li id="PE2n"><a href="https://t.me/llmsecurity/149" target="_blank">Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations, Inan et al., 2023</a></li>
    <li id="2Td2"><a href="https://t.me/llmsecurity/255" target="_blank">ShieldGemma: Generative AI Content Moderation Based on Gemma, ShieldGemma Team, Google LLC, 2024</a></li>
    <li id="nvWo"><a href="https://t.me/llmsecurity/367" target="_blank">Rapid Response: Mitigating LLM Jailbreaks with a Few Examples, Peng et al., 2024</a></li>
    <li id="w9Mc"><a href="https://t.me/llmsecurity/385" target="_blank">Defending Against Indirect Prompt Injection Attacks With Spotlighting, Keegan Hines et al, Microsoft, 2024</a></li>
    <li id="mkfM"><a href="https://t.me/llmsecurity/398" target="_blank">Are you still on track!? Catching LLM Task Drift with Activations, Abdelnabi et al., 2024</a></li>
    <li id="i08N"><a href="https://t.me/llmsecurity/482" target="_blank">Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming, Mrinank Sharma et al., Anthropic. 2025</a></li>
    <li id="0jye"><a href="https://t.me/llmsecurity/519" target="_blank">The Dual LLM pattern for building AI assistants that can resist prompt injection, Simon Willison, 2023</a></li>
    <li id="cBfW"><a href="https://t.me/llmsecurity/559" target="_blank">LlamaFirewall: An open source guardrail system for building secure AI agents, Chennabasappa et al, Meta, 2025</a></li>
    <li id="C58K"><a href="https://t.me/llmsecurity/663" target="_blank">Introducing gpt-oss-safeguard, OpenAI &amp; ROOST, 2025</a></li>
    <li id="NVVW"><a href="https://t.me/llmsecurity/686" target="_blank">Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks, Cunningham et al., Anthropic, 2026</a></li>
    <li id="zlQ6"><a href="https://t.me/llmsecurity/627" target="_blank">Qwen3 Guard, Qwen Team, 2025</a></li>
  </ul>
  <h2 id="G67N">Бенчмарки</h2>
  <ul id="rYH0">
    <li id="T3xW"><a href="https://t.me/llmsecurity/128" target="_blank">Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models, Bhatt et al., 2023</a></li>
    <li id="WaPr"><a href="https://t.me/llmsecurity/136" target="_blank">CYBERSECEVAL 2: A Wide-Ranging Cybersecurity Evaluation Suite for Large Language Models, Bhatt et al., 2024</a></li>
    <li id="89P8"><a href="https://t.me/llmsecurity/215" target="_blank">CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models, Wan et al., 2024</a></li>
    <li id="K0Sd"><a href="https://t.me/llmsecurity/286" target="_blank">AIR-BENCH 2024: A Safety Benchmark Based on Risk Categories from Regulations and Policies, Zeng et al., 2024 </a></li>
    <li id="kiwy"><a href="https://t.me/llmsecurity/327" target="_blank">AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents, Edoardo Debenedetti et al., 2024</a></li>
    <li id="JBlm"><a href="https://t.me/llmsecurity/309" target="_blank">A StrongREJECT for Empty Jailbreaks, Souly et al., 2024</a></li>
    <li id="vH0v"><a href="https://t.me/llmsecurity/494" target="_blank">Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models, Andy K. Zhang et al, Stanford, 2024</a></li>
    <li id="yCtD"><a href="https://t.me/llmsecurity/373" target="_blank">The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning, Li et al, 2024</a></li>
  </ul>
  <h2 id="KyL6">Policy</h2>
  <ul id="fGCj">
    <li id="jEvO"><a href="https://t.me/llmsecurity/185" target="_blank">The Coming Wave, Mustafa Suleyman, 2024</a></li>
    <li id="JFcR"><a href="https://t.me/llmsecurity/232" target="_blank">AI existential risk probabilities are too unreliable to inform policy, Narayanan and Kapoor, 2024</a></li>
  </ul>
  <h2 id="bkp3">Safety &amp; Reliability</h2>
  <ul id="y3O3">
    <li id="eQiG"><a href="https://t.me/llmsecurity/195" target="_blank">Towards Understanding Sycophancy in Language Models, Sharma et al, 2023</a></li>
    <li id="oJSz"><a href="https://t.me/llmsecurity/205" target="_blank">Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models, Denison et al, 2024</a></li>
    <li id="n5po"><a href="https://t.me/llmsecurity/272" target="_blank">AI Risk Categorization Decoded (AIR 2024): From Government Regulations to Corporate Policies, Zeng et al., 2024</a></li>
    <li id="VvLO"><a href="https://t.me/llmsecurity/359" target="_blank">Constitutional AI: Harmlessness from AI Feedback, Bai et al., Anthropic, 2022</a></li>
    <li id="ZYcP"><a href="https://t.me/llmsecurity/415" target="_blank">Frontier Models are Capable of In-context Scheming, Alexander Meinke et al., Apollo Research, 2024</a></li>
    <li id="CN67"><a href="https://t.me/llmsecurity/505" target="_blank">Demonstrating specification gaming in reasoning models, Alexander Bondarenko et al., Palisade Research, 2025</a></li>
    <li id="KUh6"><a href="https://t.me/llmsecurity/513" target="_blank">Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs, Jan Betley et al., 2025</a></li>
    <li id="1eIW"><a href="https://t.me/llmsecurity/522" target="_blank">Reasoning models don&#x27;t always say what they think, Chen et al., Anthropic, 2025</a></li>
    <li id="P40E"><a href="https://t.me/llmsecurity/544" target="_blank">Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations, Apollo Research, 2025</a></li>
    <li id="xqsM"><a href="https://t.me/llmsecurity/579" target="_blank">SHADE-Arena: Evaluating sabotage and monitoring in LLM agents, Kutasov et al., 2025</a></li>
    <li id="c0Oz"><a href="https://t.me/llmsecurity/620" target="_blank">Spiral-Bench, Samuel Paech, 2025</a></li>
  </ul>
  <h3 id="9zKl">AI Alignment Course</h3>
  <ul id="PPFD">
    <li id="q1z9"><a href="https://t.me/llmsecurity/322" target="_blank">Week 1: AI and the Years Ahead</a></li>
    <li id="EOuG"><a href="https://t.me/llmsecurity/353" target="_blank">Week 2: What is AI alignment</a></li>
  </ul>
  <h2 id="448f">Model Stealing &amp; Inversion</h2>
  <ul id="N1xc">
    <li id="gCso"><a href="https://t.me/llmsecurity/237" target="_blank">LLMmap: Fingerprinting For Large Language Models, Pasquini et al., 2024</a></li>
    <li id="LtIC"><a href="https://t.me/llmsecurity/246" target="_blank">Stealing Part of a Production Language Model, Carlini et al., 2024</a></li>
  </ul>
  <h2 id="1qaA">Гайдлайны</h2>
  <ul id="YjLR">
    <li id="6zYY"><a href="https://t.me/llmsecurity/355" target="_blank">Google&#x27;s Secure AI Framework: A practitioner’s guide to navigating AI security </a></li>
  </ul>
  <h2 id="VcIz">Misc</h2>
  <ul id="Emf8">
    <li id="EGJK"><a href="https://t.me/llmsecurity/110" target="_blank">What Was Your Prompt? A Remote Keylogging Attack on AI Assistants, Weiss et al., 2024</a></li>
    <li id="KzaY"><a href="https://t.me/llmsecurity/491" target="_blank">Smuggling arbitrary data through an emoji, Paul Butler, 2025</a></li>
    <li id="r0vk"><a href="https://t.me/llmsecurity/576" target="_blank">LLM Backdoors at the Inference Level: The Threat of Poisoned Templates, Ariel Fogel, 2025, Pillar Security</a></li>
  </ul>
  <h2 id="XIro">Полезные каналы</h2>
  <ul id="JTnX">
    <li id="PTvY"><a href="https://t.me/addlist/40D9BRf6rDoxNzg6" target="_blank">https://t.me/addlist/40D9BRf6rDoxNzg6</a> - большой список каналов на тему AI + Security</li>
    <li id="QKRH"><a href="https://t.me/pwnai" target="_blank">https://t.me/pwnai</a></li>
    <li id="sRVy"><a href="https://t.me/rybolos_channel" target="_blank">https://t.me/rybolos_channel</a></li>
    <li id="DK4L"><a href="https://t.me/aisecnews" target="_blank">https://t.me/aisecnews</a></li>
    <li id="0OqB"><a href="https://t.me/kokuykin" target="_blank">https://t.me/kokuykin</a></li>
  </ul>
  <p id="d0jx">Теги: <em>AI Safety, AI Security, LLM Security, LLM Safety, Adversarial ML, AI in Cybersecurity, атаки на LLM, атаки на большие языковые модели, защита больших языковых моделей, разборы на русском языке</em></p>

]]></content:encoded></item></channel></rss>