Exploring Modularity of LLM-Based Agentic Systems for Drug Discovery
Paper in proceeding, 2026

Large-language models (LLMs) incorporated into agentic systems present exciting opportunities to accelerate drug discovery. In this study, we examine the modularity of LLM-based agentic systems for drug discovery, i.e., whether parts of the system are interchangeable, a topic that has received limited attention in drug discovery. We compare the performance of different LLMs and the effectiveness of tool-calling agents versus code-generating agents. Our case study, comparing performance in orchestrating tools for chemistry and drug discovery using an LLM-as-a-judge score, shows that Claude-3.5-Sonnet, Claude-3.7-Sonnet and GPT-4o outperform alternative LLMs such as Llama-3.1-8B, Llama-3.1-70B, GPT-3.5-Turbo, and Nova-Micro. Although we confirm that code-generating agents outperform the tool-calling ones on average, we show that this is highly question- and model-dependent. Furthermore, the impact of replacing the system promptl is dependent on the question and model, underscoring that even in this particular domain one cannot replace components of the system without re-engineering. Our study highlights the necessity of further research into the modularity of LLM-based agentic systems to enable the development of reliable and modular solutions for real-world problems.

CodeAgent

LLM-as-a-judge

ToolCallingAgent

LLM-based agent modularity

drug discovery

smolagents

Author

Laura van Weesep

AstraZeneca AB

Uppsala University

Samuel Genheden

AstraZeneca AB

Ola Engkvist

Chalmers, Computer Science and Engineering (Chalmers), Data Science and AI

AstraZeneca AB

Jens Sjölund

Uppsala University

Lecture Notes in Computer Science

0302-9743 (ISSN) 1611-3349 (eISSN)

Vol. 16259 LNCS 282-298
9783032228192 (ISBN)

22nd European Conference on Multi-Agent Systems, EUMAS 2025
Bucharest, Romania,

Subject Categories (SSIF 2025)

Computer Sciences

DOI

10.1007/978-3-032-22820-8_19

More information

Latest update

9/15/2026