117Papers
4Books
1Articles
128Others
Synced from Zotero

Papers (117)

Title Author Year Description
Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits Amirhosein Ghasemabadi, Di Niu 2026 Large language models (LLMs) generate fluent and complex outputs but often fail to recognize their own mistakes and hallucinations. Existing approaches typically rely on external judges, multi-sample consistency, or text-based self-critique, which incur additional compute or correlate weakly with…
From Lagging to Leading: Validating Hard Braking Events as High-Density Indicators of Segment Crash Risk Yechen Li, Shantanu Shahane, Shoshana Vasserman et al. 2026 Identifying high crash risk road segments and accurately predicting crash incidence is fundamental to implementing effective safety countermeasures. While collision data inherently reflects risk, the infrequency and inconsistent reporting of crashes present a major challenge to robust risk…
Shaping capabilities with token-level data filtering Neil Rathi, Alec Radford 2026 Current approaches to reducing undesired capabilities in language models are largely post hoc, and can thus be easily bypassed by adversaries. A natural alternative is to shape capabilities during pretraining itself. On the proxy task of removing medical capabilities, we show that the simple…
Reinforcement Learning via Self-Distillation Jonas Hübotter, Frederike Lübeck, Lejs Behric et al. 2026 Large language models are increasingly post-trained with reinforcement learning in verifiable domains such as code and math. Yet, current methods for reinforcement learning with verifiable rewards (RLVR) learn only from a scalar outcome reward per attempt, creating a severe credit-assignment…
Learning a Generative Meta-Model of LLM Activations Grace Luo, Jiahai Feng, Trevor Darrell et al. 2026 Existing approaches for analyzing neural network activations, such as PCA and sparse autoencoders, rely on strong structural assumptions. Generative models offer an alternative: they can uncover structure without such assumptions and act as priors that improve intervention fidelity. We explore this…
LLaDA2.1: Speeding Up Text Diffusion via Token Editing Tiwei Bie, Maosong Cao, Xiang Cao et al. 2026 While LLaDA2.0 showcased the scaling potential of 100B-level block-diffusion models and their inherent parallelization, the delicate equilibrium between decoding speed and generation quality has remained an elusive frontier. Today, we unveil LLaDA2.1, a paradigm shift designed to transcend this…
Agents of Chaos Natalie Shapira, Chris Wendler, Avery Yen et al. 2026 We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under…
Discovering Multiagent Learning Algorithms with Large Language Models Zun Li, John Schultz, Daniel Hennes et al. 2026 Much of the advancement of Multi-Agent Reinforcement Learning (MARL) in imperfect-information games has historically depended on manual iterative refinement of baselines. While foundational families like Counterfactual Regret Minimization (CFR) and Policy Space Response Oracles (PSRO) rest on solid…
Attention Residuals Kimi Team, Guangyu Chen, Yu Zhang et al. 2026 Residual connections with PreNorm are standard in modern LLMs, yet they accumulate all layer outputs with fixed unit weights. This uniform aggregation causes uncontrolled hidden-state growth with depth, progressively diluting each layer's contribution. We propose Attention Residuals (AttnRes)…
Composer 2 Technical Report Cursor Research, :, Aaron Chan et al. 2026 Composer 2 is a specialized model designed for agentic software engineering. The model demonstrates strong long-term planning and coding intelligence while maintaining the ability to efficiently solve problems for interactive use. The model is trained in two phases: first, continued pretraining to…
MSA: Memory Sparse Attention for Efficient End-to-End Memory Model Scaling to 100M Tokens Yu Chen, Runkai Chen, Sheng Yi et al. 2026 Long-term memory is a cornerstone of human intelligence. Enabling AI to process lifetime-scale information remains a long-standing pursuit in the field. Due to the constraints of full-attention architectures, the effective context length of large language models (LLMs) is typically limited to 1M…
Agents of Chaos Natalie Shapira, Chris Wendler, Avery Yen et al. 2026 We report an exploratory red-teaming study of autonomous language-model-powered agents deployed in a live laboratory environment with persistent memory, email accounts, Discord access, file systems, and shell execution. Over a two-week period, twenty AI researchers interacted with the agents under…
Dive into Claude Code: The Design Space of Today's and Future AI Agent Systems Jiacheng Liu, Xiaohan Zhao, Xinyi Shang et al. 2026 Claude Code is an agentic coding tool that can run shell commands, edit files, and call external services on behalf of the user. This study describes its comprehensive architecture by analyzing the publicly available TypeScript source code and further comparing it with OpenClaw, an independent…
LeWorldModel: Stable End-to-End Joint-Embedding Predictive Architecture from Pixels Lucas Maes, Quentin Le Lidec, Damien Scieur et al. 2026 Joint Embedding Predictive Architectures (JEPAs) offer a compelling framework for learning world models in compact latent spaces, yet existing methods remain fragile, relying on complex multi-term losses, exponential moving averages, pre-trained encoders, or auxiliary supervision to avoid…
Incompressible Knowledge Probes: Estimating Black-Box LLM Parameter Counts via Factual Capacity Bojie Li 2026 Closed-source frontier labs do not disclose parameter counts, and the standard alternative -- inference economics -- carries $2\times$+ uncertainty from hardware, batching, and serving-stack assumptions external to the model. We exploit a tighter intrinsic bound: storing $F$ facts requires at least…
MAGMA: A Multi-Graph based Agentic Memory Architecture for AI Agents Dongming Jiang, Yi Li, Guanpeng Li et al. 2026 Memory-Augmented Generation (MAG) extends Large Language Models with external memory to support long-context reasoning, but existing approaches largely rely on semantic similarity over monolithic memory stores, entangling temporal, causal, and entity information. This design limits interpretability…
EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning Chuanrui Hu, Xingze Gao, Zuyi Zhou et al. 2026 Large Language Models (LLMs) are increasingly deployed as long-term interactive agents, yet their limited context windows make it difficult to sustain coherent behavior over extended interactions. Existing memory systems often store isolated records and retrieve fragments, limiting their ability to…
From Context to EDUs: Faithful and Structured Context Compression via Elementary Discourse Unit Decomposition Yiqing Zhou, Yu Lei, Shuzheng Si et al. 2026 Managing extensive context remains a critical bottleneck for Large Language Models (LLMs), particularly in applications like long-document question answering and autonomous agents where lengthy inputs incur high computational costs and introduce noise. Existing compression techniques often disrupt…
WorldMM: Dynamic Multimodal Memory Agent for Long Video Reasoning Woongyeong Yeo, Kangsan Kim, Jaehong Yoon et al. 2026 Recent advances in video large language models have demonstrated strong capabilities in understanding short clips. However, scaling them to hours- or days-long videos remains highly challenging due to limited context capacity and the loss of critical visual details during abstraction. Existing…
LightMem: Lightweight and Efficient Memory-Augmented Generation Jizhan Fang, Xinle Deng, Haoming Xu et al. 2026 Despite their remarkable capabilities, Large Language Models (LLMs) struggle to effectively leverage historical interaction information in dynamic and complex environments. Memory systems enable LLMs to move beyond stateless interactions by introducing persistent information storage, retrieval, and…
RGMem: Renormalization Group-inspired Memory Evolution for Language Agents Ao Tian, Yunfeng Lu, Xinxin Fan et al. 2026 Personalized and continuous interactions are critical for LLM-based conversational agents, yet finite context windows and static parametric memory hinder the modeling of long-term, cross-session user states. Existing approaches, including retrieval-augmented generation and explicit memory systems…
What Deserves Memory: Adaptive Memory Distillation for LLM Agents Wenquan Ma, Jiayan Nan, Wenlong Wu et al. 2026 Memory systems for LLM agents struggle to determine what information deserves retention. Existing approaches rely on predefined heuristics such as importance scores, emotional tags, or factual templates, encoding designer intuition rather than learning from the data itself. Inspired by cognitive…
Remember Me, Refine Me: A Dynamic Procedural Memory Framework for Experience-Driven Agent Evolution Zouying Cao, Jiaji Deng, Li Yu et al. 2026 Procedural memory enables large language model (LLM) agents to internalize "how-to" knowledge, theoretically reducing redundant trial-and-error. However, existing frameworks predominantly suffer from a "passive accumulation" paradigm, treating memory as a static append-only archive. To bridge the…
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models Qizheng Zhang, Changran Hu, Shubhangi Upasani et al. 2026 Large language model (LLM) applications such as agents and domain-specific reasoning increasingly rely on context adaptation: modifying inputs with instructions, strategies, or evidence, rather than weight updates. Prior approaches improve usability but often suffer from brevity bias, which drops…
Memory in the Age of AI Agents Yuyang Hu, Shichun Liu, Yanwei Yue et al. 2026 Memory has emerged, and will continue to remain, a core capability of foundation model-based agents. As research on agent memory rapidly expands and attracts unprecedented attention, the field has also become increasingly fragmented. Existing works that fall under the umbrella of agent memory often…
Agent Harness for Large Language Model Agents: A Survey Qianyu Meng, Yanan Wang, Liyi Chen et al. 2026 Nowadays, the reliability of large language model (LLM) agents in production environments is increasingly determined not by the underlying model but by the agent harness that encapsulates it. As tasks grow longer and more complex, recent studies demonstrate order-of-magnitude reliability gains…
Skillful joint probabilistic weather forecasting from marginals Ferran Alet, Ilan Price, Andrew El-Kadi et al. 2025 Machine learning (ML)-based weather models have rapidly risen to prominence due to their greater accuracy and speed than traditional forecasts based on numerical weather prediction (NWP), recently outperforming traditional ensembles in global probabilistic weather forecasting. This paper presents…
AI and Generative AI Transforming Disaster Management: A Survey of Damage Assessment and Response Techniques Aman Raj, Ankit Shetgaonkar, Lakshit Arora et al. 2025 Natural disasters, including earthquakes, wildfires and cyclones, bear a huge risk on human lives as well as infrastructure assets. An effective response to disaster depends on the ability to rapidly and efficiently assess the intensity of damage. Artificial Intelligence (AI) and Generative…
Confucius Code Agent: Scalable Agent Scaffolding for Real-World Codebases Sherman Wong, Zhenting Qi, Zhaodong Wang et al. 2025 Real-world software engineering tasks require coding agents that can operate over massive repositories, sustain long-horizon sessions, and reliably coordinate complex toolchains at test time. Existing research-grade coding agents offer transparency but struggle when scaled to heavier…
Recursive Language Models Alex L. Zhang, Tim Kraska, Omar Khattab 2025 We study allowing large language models (LLMs) to process arbitrarily long prompts through the lens of inference-time scaling. We propose Recursive Language Models (RLMs), a general inference strategy that treats long prompts as part of an external environment and allows the LLM to programmatically…
Aligning machine and human visual representations across abstraction levels Lukas Muttenthaler, Klaus Greff, Frieda Born et al. 2025 Abstract Deep neural networks have achieved success across a wide range of applications, including as models of human behaviour and neural representations in vision tasks 1,2 . However, neural network training and human learning differ in fundamental ways, and neural networks often fail to…
PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations Ruining He, Lukasz Heldt, Lichan Hong et al. 2025 Large Language Models (LLMs) pose a new paradigm of modeling and computation for information tasks. Recommendation systems are a critical application domain poised to benefit significantly from the sequence modeling capabilities and world knowledge inherent in these large models. In this paper, we…
PaperBench: Evaluating AI's Ability to Replicate AI Research Giulio Starace, Oliver Jaffe, Dane Sherburn et al. 2025 We introduce PaperBench, a benchmark evaluating the ability of AI agents to replicate state-of-the-art AI research. Agents must replicate 20 ICML 2024 Spotlight and Oral papers from scratch, including understanding paper contributions, developing a codebase, and successfully executing experiments…
Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training Junkai Zhang, Zihao Wang, Lin Gui et al. 2025 Reinforcement fine-tuning (RFT) often suffers from \emph{reward over-optimization}, where a policy model hacks the reward signals to achieve high scores while producing low-quality outputs. Our theoretical analysis shows that the key lies in reward misspecification at the high-reward tail: the…
Text-to-LoRA: Instant Transformer Adaption Rujikorn Charakorn, Edoardo Cetin, Yujin Tang et al. 2025 While Foundation Models provide a general tool for rapid content creation, they regularly require task-specific adaptation. Traditionally, this exercise involves careful curation of datasets and repeated fine-tuning of the underlying model. Fine-tuning techniques enable practitioners to adapt…
From Language to Action: A Review of Large Language Models as Autonomous Agents and Tool Users Sadia Sultana Chowa, Riasad Alvi, Subhey Sadi Rahman et al. 2025 The pursuit of human-level artificial intelligence (AI) has significantly advanced the development of autonomous agents and Large Language Models (LLMs). LLMs are now widely utilized as decision-making agents for their ability to interpret instructions, manage sequential tasks, and adapt through…
MemVerse: Multimodal Memory for Lifelong Learning Agents Junming Liu, Yifei Sun, Weihua Cheng et al. 2025 Despite rapid progress in large-scale language and vision models, AI agents still suffer from a fundamental limitation: they cannot remember. Without reliable memory, agents catastrophically forget past experiences, struggle with long-horizon reasoning, and fail to operate coherently in multimodal…
MMAG: Mixed Memory-Augmented Generation for Large Language Models Applications Stefano Zeppieri 2025 Large Language Models (LLMs) excel at generating coherent text within a single prompt but fall short in sustaining relevance, personalization, and continuity across extended interactions. Human communication, however, relies on multiple forms of memory, from recalling past conversations to adapting…
Sophia: A Persistent Agent Framework of Artificial Life Mingyang Sun, Feng Hong, Weinan Zhang 2025 The development of LLMs has elevated AI agents from task-specific tools to long-lived, decision-making entities. Yet, most architectures remain static and reactive, tethered to manually defined, narrow scenarios. These systems excel at perception (System 1) and deliberation (System 2) but lack a…
Memoria: A Scalable Agentic Memory Framework for Personalized Conversational AI Samarth Sarin, Lovepreet Singh, Bhaskarjit Sarmah et al. 2025 Agentic memory is emerging as a key enabler for large language models (LLM) to maintain continuity, personalization, and long-term context in extended user interactions, critical capabilities for deploying LLMs as truly interactive and adaptive agents. Agentic memory refers to the memory that…
A Simple Yet Strong Baseline for Long-Term Conversational Memory of LLM Agents Sizhe Zhou, Jiawei Han 2025 LLM-based conversational agents still struggle to maintain coherent, personalized interaction over many sessions: fixed context windows limit how much history can be kept in view, and most external memory approaches trade off between coarse retrieval over large chunks and fine-grained but…
General Agentic Memory Via Deep Research B. Y. Yan, Chaofan Li, Hongjin Qian et al. 2025 Memory is critical for AI agents, yet the widely-adopted static memory, aiming to create readily available memory in advance, is inevitably subject to severe information loss. To address this limitation, we propose a novel framework called \textbf{general agentic memory (GAM)}. GAM follows the…
O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents Piaohong Wang, Motong Tian, Jiaxian Li et al. 2025 Recent advancements in LLM-powered agents have demonstrated significant potential in generating human-like responses; however, they continue to face challenges in maintaining long-term interactions within complex environments, primarily due to limitations in contextual consistency and dynamic…
RCR-Router: Efficient Role-Aware Context Routing for Multi-Agent LLM Systems with Structured Memory Jun Liu, Zhenglun Kong, Changdi Yang et al. 2025 Multi-agent large language model (LLM) systems have shown strong potential in complex reasoning and collaborative decision-making tasks. However, most existing coordination schemes rely on static or full-context routing strategies, which lead to excessive token consumption, redundant memory…
Enabling Personalized Long-term Interactions in LLM-based Agents through Persistent Memory and User Profiles Rebecca Westhäußer, Wolfgang Minker, Sebatian Zepf 2025 Large language models (LLMs) increasingly serve as the central control unit of AI agents, yet current approaches remain limited in their ability to deliver personalized interactions. While Retrieval Augmented Generation enhances LLM capabilities by improving context-awareness, it lacks mechanisms…
Livia: An Emotion-Aware AR Companion Powered by Modular AI Agents and Progressive Memory Compression Rui Xi, Xianghan Wang 2025 Loneliness and social isolation pose significant emotional and health challenges, prompting the development of technology-based solutions for companionship and emotional support. This paper introduces Livia, an emotion-aware augmented reality (AR) companion app designed to provide personalized…
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree Xiang Lei, Qin Li, Min Zhang et al. 2025 Large Language Models (LLMs) often exhibit factual inconsistencies and logical decay in extended, multi-turn dialogues, a challenge stemming from their reliance on static, pre-trained knowledge and an inability to reason adaptively over the dialogue history. Prevailing mitigation strategies, such…
WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research Zijian Li, Xin Guan, Bo Zhang et al. 2025 This paper tackles \textbf{open-ended deep research (OEDR)}, a complex challenge where AI agents must synthesize vast web-scale information into insightful reports. Current approaches are plagued by dual-fold limitations: static research pipelines that decouple planning from evidence acquisition…
CAM: A Constructivist View of Agentic Memory for LLM-Based Reading Comprehension Rui Li, Zeyu Zhang, Xiaohe Bo et al. 2025 Current Large Language Models (LLMs) are confronted with overwhelming information volume when comprehending long-form documents. This challenge raises the imperative of a cohesive memory module, which can elevate vanilla LLMs into autonomous reading agents. Despite the emergence of some heuristic…
Pre-Storage Reasoning for Episodic Memory: Shifting Inference Burden to Memory for Personalized Dialogue Sangyeop Kim, Yohan Lee, Sanghwa Kim et al. 2025 Effective long-term memory in conversational AI requires synthesizing information across multiple sessions. However, current systems place excessive reasoning burden on response generation, making performance significantly dependent on model sizes. We introduce PREMem (Pre-storage Reasoning for…
Mem-α: Learning Memory Construction via Reinforcement Learning Yu Wang, Ryuichi Takanobu, Zhiqi Liang et al. 2025 Large language model (LLM) agents are constrained by limited context windows, necessitating external memory systems for long-term information understanding. Current memory-augmented agents typically depend on pre-defined instructions and tools for memory updates. However, language models may lack…
SGMem: Sentence Graph Memory for Long-Term Conversational Agents Yaxiong Wu, Yongyue Zhang, Sheng Liang et al. 2025 Long-term conversational agents require effective memory management to handle dialogue histories that exceed the context window of large language models (LLMs). Existing methods based on fact extraction or summarization reduce redundancy but struggle to organize and retrieve relevant information…
In Prospect and Retrospect: Reflective Memory Management for Long-term Personalized Dialogue Agents Zhen Tan, Jun Yan, I-Hung Hsu et al. 2025 Large Language Models (LLMs) have made significant progress in open-ended dialogue, yet their inability to retain and retrieve relevant information from long-term interactions limits their effectiveness in applications requiring sustained personalization. External memory mechanisms have been…
Recommender AI Agent: Integrating Large Language Models for Interactive Recommendations Xu Huang, Jianxun Lian, Yuxuan Lei et al. 2025 Recommender models capture ever-changing user preferences by training with in-domain user behavior data. These models are typically lightweight, facilitating real-time and large-scale online services. However, these models often falter when tasked with providing more sophisticated functionalities…
ELDER: Enhancing Lifelong Model Editing with Mixture-of-LoRA Jiaang Li, Quan Wang, Zhongnan Wang et al. 2025 Large language models (LLMs) require model editing to efficiently update specific knowledge within them and avoid factual errors. Most model editing methods are solely designed for single-time use and result in a significant forgetting effect in lifelong editing scenarios, where sequential edits…
Hindsight is 20/20: Building Agent Memory that Retains, Recalls, and Reflects Chris Latimer, Nicoló Boschi, Andrew Neeser et al. 2025 Agent memory has been touted as a dimension of growth for LLM-based applications, enabling agents that can accumulate experience, adapt across sessions, and move beyond single-shot question answering. The current generation of agent memory systems treats memory as an external layer that extracts…
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience Zeyi Sun, Ziyu Liu, Yuhang Zang et al. 2025 Repurposing large vision-language models (LVLMs) as computer use agents (CUAs) has led to substantial breakthroughs, primarily driven by human-labeled data. However, these models often struggle with novel and specialized software, particularly in scenarios lacking human annotations. To address this…
Agent Workflow Memory Zora Zhiruo Wang, Jiayuan Mao, Daniel Fried et al. 2025 Despite the potential of language model-based agents to solve real-world tasks such as web navigation, current methods still struggle with long-horizon tasks with complex action trajectories. In contrast, humans can flexibly solve complex tasks by learning reusable task workflows from past…
JARVIS-1: Open-World Multi-Task Agents With Memory-Augmented Multimodal Language Models Zihao Wang, Shaofei Cai, Anji Liu et al. 2025 Achieving human-like planning and control with multimodal observations in an open world is a key milestone for more functional generalist agents. Existing approaches can handle certain long-horizon tasks in an open world. However, they still struggle when the number of open-world tasks could…
MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation Hongjin Qian, Zheng Liu, Peitian Zhang et al. 2025
SoftCoT: Soft Chain-of-Thought for Efficient Reasoning with LLMs Yige Xu, Xu Guo, Zhiwei Zeng et al. 2025 Chain-of-Thought (CoT) reasoning enables Large Language Models (LLMs) to solve complex reasoning tasks by generating intermediate reasoning steps. However, most existing approaches focus on hard token decoding, which constrains reasoning within the discrete vocabulary space and may not always be…
SeCom: On Memory Construction and Retrieval for Personalized Conversational Agents Zhuoshi Pan, Qianhui Wu, Huiqiang Jiang et al. 2024 To deliver coherent and personalized experiences in long-term conversations, existing approaches typically perform retrieval augmented response generation by constructing memory banks from conversation history at either the turn-level, session-level, or through summarization techniques. In this…
Memolet: Reifying the Reuse of User-AI Conversational Memories Ryan Yen, Jian Zhao 2024
Human-inspired Episodic Memory for Infinite Context LLMs Zafeirios Fountas, Martin Benfeghoul, Adnan Oomerjee et al. 2024 Large language models (LLMs) have shown remarkable capabilities, but still struggle with processing extensive contexts, limiting their ability to maintain coherence and accuracy over long sequences. In contrast, the human brain excels at organising and retrieving episodic experiences across vast…
Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation Wazeer Deen Zulfikar, Samantha Chan, Pattie Maes 2024
RoleLLM: Benchmarking, Eliciting, and Enhancing Role-Playing Abilities of Large Language Models Noah Wang, Z.y. Peng, Haoran Que et al. 2024 The advent of Large Language Models (LLMs) has paved the way for complex tasks such as role-playing, which enhances user interactions by enabling models to imitate various characters. However, the closed-source nature of state-of-the-art LLMs and their general-purpose training limit role-playing…
MovieChat: From Dense Token to Sparse Memory for Long Video Understanding Enxin Song, Wenhao Chai, Guanhong Wang et al. 2024 Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory…
Self-Updatable Large Language Models by Integrating Context into Model Parameters Yu Wang, Xinshuang Liu, Xiusi Chen et al. 2024 Despite significant advancements in large language models (LLMs), the rapid and frequent integration of small-scale experiences, such as interactions with sur- rounding objects, remains a substantial challenge. Two critical factors in assimilating these experiences are (1) **Efficacy**: the ability…
WISE: Rethinking the Knowledge Memory for Lifelong Model Editing of Large Language Models Peng Wang, Zexi Li, Ningyu Zhang et al. 2024
Online Adaptation of Language Models with a Memory of Amortized Contexts Jihoon Tack, Jaehyung Kim, Eric Mitchell et al. 2024
Neighboring Perturbations of Knowledge Editing on Large Language Models Jun-Yu Ma, Zhen-Hua Ling, Ningyu Zhang et al. 2024 Despite their exceptional capabilities, large language models (LLMs) are prone to generating unintended text due to false or outdated knowledge. Given the resource-intensive nature of retraining LLMs, there has been a notable increase in the development of knowledge editing. However, current…
CharacterGLM: Customizing Social Characters with Large Language Models Jinfeng Zhou, Zhuang Chen, Dazhen Wan et al. 2024 Character-based dialogue (CharacterDial) has become essential in the industry (e.g., Character.AI), enabling users to freely customize social characters for social interactions. However, the generalizability and adaptability across various conversational scenarios inherent in customizing social…
SAGE: Self-evolving Agents with Reflective and Memory-augmented Abilities 2024
Buffer of Thoughts: Thought-Augmented Reasoning with Large Language Models Ling Yang, Zhaochen Yu, Tianjun Zhang et al. 2024
RecMind: Large Language Model Powered Agent For Recommendation Yancheng Wang, Ziyan Jiang, Zheng Chen et al. 2024 While the recommendation system (RS) has advanced significantly through deep learning, current RS approaches usually train and fine-tune models on task-specific datasets, limiting their generalizability to new recommendation tasks and their ability to leverage external knowledge due to model scale…
ExpeL: LLM Agents Are Experiential Learners Andrew Zhao, Daniel Huang, Quentin Xu et al. 2024 The recent surge in research interest in applying large language models (LLMs) to decision-making tasks has flourished by leveraging the extensive world knowledge embedded in LLMs. While there is a growing demand to tailor LLMs for custom decision-making tasks, finetuning them for specific tasks is…
RazorAttention: Efficient KV Cache Compression Through Retrieval Heads Hanlin Tang, Yang Lin, Jing Lin et al. 2024 The memory and computational demands of Key-Value (KV) cache present significant challenges for deploying long-context language models. Previous approaches attempt to mitigate this issue by selectively dropping tokens, which irreversibly erases critical information that might be needed for future…
SnapKV: LLM Knows What You are Looking for Before Generation Yuhong Li, Yingbing Huang, Bowen Yang et al. 2024
Taking a Deep Breath: Enhancing Language Modeling of Large Language Models with Sentinel Tokens Weiyao Luo, Suncong Zheng, Heming Xia et al. 2024 Large language models (LLMs) have shown promising efficacy across various tasks, becoming powerful tools in numerous aspects of human life. However, Transformer-based LLMs suffer a performance degradation when modeling long-term contexts due to they discard some information to reduce computational…
Were RNNs All We Needed? Leo Feng, Frederick Tung, Mohamed Osama Ahmed et al. 2024 The introduction of Transformers in 2017 reshaped the landscape of deep learning. Originally proposed for sequence modelling, Transformers have since achieved widespread success across various domains. However, the scalability limitations of Transformers - particularly with respect to sequence…
Attention Is All You Need Ashish Vaswani, Noam Shazeer, Niki Parmar et al. 2023 The dominant sequence transduction models are based on complex recurrent or convolutional neural networks in an encoder-decoder configuration. The best performing models also connect the encoder and decoder through an attention mechanism. We propose a new simple network architecture, the…
CALYPSO: LLMs as Dungeon Master's Assistants Andrew Zhu, Lara Martin, Andrew Head et al. 2023 The role of a Dungeon Master, or DM, in the game Dungeons & Dragons is to perform multiple tasks simultaneously. The DM must digest information about the game setting and monsters, synthesize scenes to present to other players, and respond to the players' interactions with the scene. Doing all of…
Prompted LLMs as Chatbot Modules for Long Open-domain Conversation Gibbeum Lee, Volker Hartmann, Jongho Park et al. 2023 In this paper, we propose MPC (Modular Prompted Chatbot), a new approach for creating high-quality conversational agents without the need for fine-tuning. Our method utilizes pre-trained large language models (LLMs) as individual modules for long-term consistency and flexibility, by using…
MemoryBank: Enhancing Large Language Models with Long-Term Memory Wanjun Zhong, Lianghong Guo, Qiqi Gao et al. 2023 Revolutionary advancements in Large Language Models have drastically reshaped our interactions with artificial intelligence systems. Despite this, a notable hindrance remains-the deficiency of a long-term memory mechanism within these models. This shortfall becomes increasingly evident in…
Character-LLM: A Trainable Agent for Role-Playing Yunfan Shao, Linyang Li, Junqi Dai et al. 2023 Large language models (LLMs) can be used to serve as agents to simulate human behaviors, given the powerful ability to understand human instructions and provide high-quality generated texts. Such ability stimulates us to wonder whether LLMs can simulate a person in a higher form than simple human…
Efficient Episodic Memory Utilization of Cooperative Multi-Agent Reinforcement Learning Hyungho Na, Yunkyeong Seo, Il-chul Moon 2023 In cooperative multi-agent reinforcement learning (MARL), agents aim to achieve a common goal, such as defeating enemies or scoring a goal. Existing MARL algorithms are effective but still require significant learning time and often get trapped in local optima by complex tasks, subsequently failing…
CREATOR: Tool Creation for Disentangling Abstract and Concrete Reasoning of Large Language Models Cheng Qian, Chi Han, Yi Fung et al. 2023 Large Language Models (LLMs) have made significant progress in utilizing tools, but their ability is limited by API availability and the instability of implicit reasoning, particularly when both planning and execution are involved. To overcome these limitations, we propose CREATOR, a novel…
Reflexion: Language Agents with Verbal Reinforcement Learning Noah Shinn, Federico Cassano, Edward Berman et al. 2023 Large language models (LLMs) have been increasingly used to interact with external environments (e.g., games, compilers, APIs) as goal-driven agents. However, it remains challenging for these language agents to quickly and efficiently learn from trial-and-error as traditional reinforcement learning…
A Machine with Short-Term, Episodic, and Semantic Memory Systems Taewoon Kim, Michael Cochez, Vincent Francois-Lavet et al. 2023 Inspired by the cognitive science theory of the explicit human memory systems, we have modeled an agent with short-term, episodic, and semantic memory systems, each of which is modeled with a knowledge graph. To evaluate this system and analyze the behavior of this agent, we designed and released…
Efficient Streaming Language Models with Attention Sinks Guangxuan Xiao, Yuandong Tian, Beidi Chen et al. 2023 Deploying Large Language Models (LLMs) in streaming applications such as multi-round dialogue, where long interactions are expected, is urgently needed but poses two major challenges. Firstly, during the decoding stage, caching previous tokens' Key and Value states (KV) consumes extensive memory…
Augmenting Language Models with Long-Term Memory Weizhi Wang, Li Dong, Hao Cheng et al. 2023
Adapting Language Models to Compress Contexts Alexis Chevalier, Alexander Wettig, Anirudh Ajith et al. 2023
Learning to Compress Prompts with Gist Tokens Jesse Mu, Xiang Li, Noah Goodman 2023
Scissorhands: Exploiting the Persistence of Importance Hypothesis for LLM KV Cache Compression at Test Time Zichang Liu, Aditya Desai, Fangshuo Liao et al. 2023
Focused Transformer: Contrastive Training for Context Scaling Szymon Tworkowski, Konrad Staniszewski, Mikołaj Pacek et al. 2023
H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models Zhenyu Zhang, Ying Sheng, Tianyi Zhou et al. 2023
Fast Model Editing at Scale Eric Mitchell, Charles Lin, Antoine Bosselut et al. 2021 While large pre-trained models have enabled impressive results on a variety of downstream tasks, the largest existing models still make errors, and even accurate predictions may become outdated over time. Because detecting all such failures at training time is impossible, enabling both developers…
K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters Ruize Wang, Duyu Tang, Nan Duan et al. 2021
Detecting Local Insights from Global Labels: Supervised and Zero-Shot Sequence Labeling via a Convolutional Decomposition Allen Schmaltz 2021 Abstract We propose a new, more actionable view of neural network interpretability and data analysis by leveraging the remarkable matching effectiveness of representations derived from deep networks, guided by an approach for class-conditional feature detection. The decomposition of the…
Memorizing Transformers Yuhuai Wu, Markus Norman Rabe, DeLesley Hutchins et al. 2021 Language models typically need to be trained or finetuned in order to acquire new knowledge, which involves updating their weights. We instead envision language models that can simply read and memorize new data at inference time, thus acquiring new knowledge immediately. In this work, we extend…
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism Yanping Huang, Youlong Cheng, Ankur Bapna et al. 2019 Scaling up deep neural network capacity has been known as an effective approach to improving model quality for several different machine learning tasks. In many cases, increasing model capacity beyond the memory limit of a single accelerator has required developing special algorithms or…
Pointer Networks Oriol Vinyals, Meire Fortunato, Navdeep Jaitly 2017 We introduce a new neural architecture to learn the conditional probability of an output sequence with elements that are discrete tokens corresponding to positions in an input sequence. Such problems cannot be trivially addressed by existent approaches such as sequence-to-sequence and Neural Turing…
Neural Message Passing for Quantum Chemistry Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley et al. 2017 Supervised learning on molecules has incredible potential to be useful in chemistry, drug discovery, and materials science. Luckily, several promising and closely related neural network models invariant to molecular symmetries have already been described in the literature. These models learn a…
A simple neural network module for relational reasoning Adam Santoro, David Raposo, David G. T. Barrett et al. 2017 Relational reasoning is a central component of generally intelligent behavior, but has proven difficult for neural networks to learn. In this paper we describe how to use Relation Networks (RNs) as a simple plug-and-play module to solve problems that fundamentally hinge on relational reasoning. We…
AdaGAN: Boosting Generative Models Ilya Tolstikhin, Sylvain Gelly, Olivier Bousquet et al. 2017 Generative Adversarial Networks (GAN) (Goodfellow et al., 2014) are an effective method for training generative models of complex data such as natural images. However, they are notoriously hard to train and can suffer from the problem of missing modes where the model is not able to produce examples…
Order Matters: Sequence to sequence for sets Oriol Vinyals, Samy Bengio, Manjunath Kudlur 2016 Sequences have become first class citizens in supervised learning thanks to the resurgence of recurrent neural networks. Many complex tasks that require mapping from or to a sequence of observations can now be formulated with the sequence-to-sequence (seq2seq) framework which employs the chain rule…
Multi-Scale Context Aggregation by Dilated Convolutions Fisher Yu, Vladlen Koltun 2016 State-of-the-art models for semantic segmentation are based on adaptations of convolutional networks that had originally been designed for image classification. However, dense prediction and image classification are structurally different. In this work, we develop a new convolutional network module…
Neural Machine Translation by Jointly Learning to Align and Translate Dzmitry Bahdanau, Kyunghyun Cho, Yoshua Bengio 2016 Neural machine translation is a recently proposed approach to machine translation. Unlike the traditional statistical machine translation, the neural machine translation aims at building a single neural network that can be jointly tuned to maximize the translation performance. The models proposed…
Identity Mappings in Deep Residual Networks Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. 2016 Deep residual networks have emerged as a family of extremely deep architectures showing compelling accuracy and nice convergence behaviors. In this paper, we analyze the propagation formulations behind the residual building blocks, which suggest that the forward and backward signals can be directly…
Recurrent Neural Network Regularization Wojciech Zaremba, Ilya Sutskever, Oriol Vinyals 2015 We present a simple regularization technique for Recurrent Neural Networks (RNNs) with Long Short-Term Memory (LSTM) units. Dropout, the most successful technique for regularizing neural networks, does not work well with RNNs and LSTMs. In this paper, we show how to correctly apply dropout to…
Deep Residual Learning for Image Recognition Kaiming He, Xiangyu Zhang, Shaoqing Ren et al. 2015 Deeper neural networks are more difficult to train. We present a residual learning framework to ease the training of networks that are substantially deeper than those used previously. We explicitly reformulate the layers as learning residual functions with reference to the layer inputs, instead of…
ELLA: An Efficient Lifelong Learning Algorithm Paul Ruvolo, Eric Eaton 2013 The problem of learning multiple consecutive tasks, known as lifelong learning, is of great importance to the creation of intelligent, general-purpose, and flexible machines. In this paper, we develop a method for online multi-task learning in the lifelong learning setting. The proposed Efficient…
ImageNet Classification with Deep Convolutional Neural Networks Alex Krizhevsky, Ilya Sutskever, Geoffrey E Hinton 2012
Nested Learning: The Illusion of Deep Learning Architecture Ali Behrouz, Meisam Razaviyayn, Peilin Zhong et al. Over the last decades, developing more powerful neural architectures and simultaneously designing optimization algorithms to effectively train them have been the core of research efforts to enhance the capability of machine learning models. Despite the recent progresses, particularly in developing…
Learning Domain-Driven Design Vlad Khononov
Authors Maxim Massenkoff and Peter McCrory Ruth Appel, Tim Belonax, Keir Bradwell et al.
On Layer Normalization in the Transformer Architecture Ruibin Xiong, Yunchang Yang, Di He et al. The Transformer is widely used in natural language processing tasks. To train a Transformer however, one usually needs a carefully designed learning rate warm-up stage, which is shown to be crucial to the final performance but will slow down the optimization and bring more hyperparameter tunings. In…
Page 1 of 1

Articles (1)

Title Author Year Description
Rubric-Based Rewards for RL Deep (Learning) Focus 2026 Extending the benefits of large-scale RL training to non-verifiable domains...
Page 1 of 1