1
WEB-PDE-LLM: A Multi-Agent Framework for Web-Based PDE Solving
Xiaodong An (Georgia Tech)
Guangze Luo (Scale AI)
Prof. Flavio H. Fenton (Georgia Tech)
2
1. Introduction: LLM-Aided PDE Solving
2. Methodology: The WEB-PDE-LLM Framework
3. Experimental Setup: Datasets and Metrics
4. Results and Ablation Studies
5. Discussion and Future Work
Outline
3
Designing accurate and efficient PDE solver still remains as a challenge for non-expert
Introduction: LLM for Partial Differential Equation
Computing Env
Boundary Condition
Solver Coding
Solver Debug
…
4
Introduction: Recent Success of Large Language Model (LLM)
LLMs continue to demonstrate better capabilities in code generation and scientific programming.
So what about using LLM on PDE solving?
5
Introduction: LLM-Aided PDE Solving
Current state of the arts (SOTAs) for LLM-Aided PDE solving frameworks:
CodePDE (Li et al., 2025)
PINNsAgent (Wuwu et al., 2025)
MCP-SIM (Park et al., 2026)
OP-Inf LLM (Wang et al., 2026)
Solver.py
PDE
Multi-Agent Loop
Solver.py (FEniCS)
PDE
Multi-Agent Loop
PINN.py
PDE
Multi-Agent Loop
Solver
Train
OpInf.py
PDE
Multi-Agent Loop
Solver
Train
6
Introduction: LLM-Aided PDE Solving
While these SOTA methods represent significant progress, they face the following limitations:
Cross-Platform Compatibility | | MCP-SIM only runs on linux |
No Environment Setup | | All of them requires certain env such as pyTorch and FEniCS |
Easy Visualization | | All of them struggle with real-time visualization of the PDE solving. |
Self-Debugging | | Neither PINNsAgent nor Op-Inf LLM can self-debug to decrease errors. |
Zero Pre-Training Required | | PINNsAgent and Op-Inf LLM require pre training. |
7
Method: WEB-PDE-LLM
We propose WEB-PDE-LLM, a multi-agent framework that translates natural language descriptions into browser-executable WebGL/GLSL solvers.
By calculating directly on the browser's GPU, it eliminates the need for Python environments. This zero-configuration pipeline not only guarantees cross-platform compatibility across modern devices but also enables users to visualize complex numerical simulation processes in real time.
8
Method: WEB-PDE-LLM
9
Method: WEB-PDE-LLM
10
Method: WEB-PDE-LLM (Parse Agent)
11
Method: WEB-PDE-LLM (Code Agent)
12
Method: WEB-PDE-LLM (Debug Agent)
13
Evaluation: Metric
Upon successfully generating the solver and executing the simulation, we evaluate its performance using metrics adapted from the PDEAgent-Bench (Hang et al., 2026), specifically focusing on normalized Root Mean Square Error (nRMSE), pass rate, Solvetime, and total API token cost.
We evaluate LLM-aided PDE solving by running 10 trials per PDE, recording the success rate based on successful compilation and a low nRMSE.
The time required for the LLM-generated solver to simulate the PDE up to t_end.
LLM API cost for running experiments.
14
Evaluation: PDE of Interest (5 PDEs)
Our evaluation benchmark includes five PDEs:: the advection, Burgers', and heat equations, along with the complex Fenton-Karma and ten Tusscher–Noble–Noble–Panfilov (TNNP) cardiac equations.
Advection/Burger’s (1D)
Heat (2D)
Reaction Diffusion
Fenton-Karma/TNNP (2D)
15
Evaluation: LLM of Interest
Our evaluation benchmark includes two LLMs and two Small Language Models (SLMs): GPT-5.4, Gemini-3.5-flash, Claude Opus 4.8, deepseek-v4-flash and Qwen-3.7
GPT5.4
Gemini-3.5-flash
Claude Opus 4.8
deepseek-v4-flash
Qwen-3.7
Current Scope
Future Scope
16
Main Result (nRMSE)
Op-Inf LLM is excluded because its prompts only work for its own predefined PDEs.
17
Main Result (Pass Rate)
18
Main Result (Token Cost)
19
Main Result (Solve Time)
The time required for the LLM-generated solver to simulate the PDE up to t_end.
20
Main Result (Ablation Test)
21
Main Result (Score Chart)
22
Result Visualization (by Gemini-3.5-Flash)
Advection
Burgers
Heat
Fenton -Karma
TNNP
23
Summary
We propose WEB-PDE-LLM: A new framework that generates PDE solvers that run directly in the browser.
Zero environment setup: It runs on any modern device with a standard web browser—no installation required.
Outperforms SOTA: Delivers better nRMSE and pass rates than current SOTAs, especially on complex PDEs like TNNP.
Highly cost-effective: The running cost is affordable, remaining similar to vanilla LLM API cost.
24
Future Steps
Extend evaluation to additional LLMs, including DeepSeek and Qwen.
Test more on small language models, such as GPT-OSS-20B and Gemma 4 31B.
Incorporate more complex PDE benchmarks, such as OVVR.
25
Acknowledgements
Thanks to my advisor Dr. Flavio for his great support.
Thanks to Guangze Luo for help with the LLM benchmarking method.
This work was supported by GR00027195.
26
Questions?