PhySimCode

A Benchmark and Evaluation Method for
Physics Video → Code Generation

1Machine Learning Department, CMU 2Computer Science Department, CMU 3Robotics Institute, CMU 4Electrical & Computer Engineering, CMU 5Amazon AGI
Carnegie Mellon University AmazonAGI

✉ Corresponding author & project lead: sourajak@cs.cmu.edu

TL;DR

Can a multimodal LLM watch a 2D or 3D physics simulation, work out the involved physics laws and parameter values, and write code from scratch that regenerates it?

PhySimCode is a benchmark of 160.6K procedurally generated physics videos covering 162 phenomena, rendered by two different engines (SciPy in 2D, PyBullet in 3D). Every video comes with the code that made it, its true parameter values and a step-by-step physics explanation, so a model's answer can be checked at every level.

The taskFrom pixels to physics to programs

1Inputa physics video + engine name → 2MLLMreasons about the physics → 3Outputphysical law · parameter values · Python code → 4Run itregenerated video, scored vs. input
MLLM
cot.jsonsimulation.pyterminal

          
Regenerated video from the model's code
waiting for code…
160,614video–code–CoT–parameter tuples
162physics phenomena
9physical domains
154named physical laws
engine-agnosticdata generation pipeline

Key findingsThree consistent gaps

4.5/5 vs 2.47/5

Knowing ≠ inferring

Most models state physical laws correctly (7 of 10 score >4.5/5), but the best (GPT-5 Mini) scores only 2.47/5 at recognizing the law demonstrated in the video.

7.7%

Parameters are rarely recovered

The best model (Claude Sonnet 4.6) names and estimates within ±20% only 7.7% of all ground-truth parameters.

κ 0.72 vs 0.20

Pixels ≠ dynamics

DINOv2 and VideoCLIP agree with each other (κ = 0.717) but barely with human judgments of physics equivalence (κ ≤ 0.205). Models match appearance and miss dynamics.

Benchmark resultsTen MLLMs on 2,430 videos

162 experiments × 15 uniformly random samples (seed 42). Law scores are averaged over three LLM judges (Gemini-2.5-Flash, Grok-4.1-Fast, GPT-5-Nano). Human scores are the mean of four expert raters over 100 random samples per model. Hover a chart for exact values; click legend entries to toggle them.

What we measure

👁

(i) Video understanding

Extract visual features and track objects across frames.
Measured by human appearance equivalence and DINOv2 / VideoCLIP similarity.

∑

(ii) Physics understanding

Bind what it sees to what it knows: identify the governing law and estimate parameters.
Measured by law correctness vs. equivalence and parameter recovery (±20%).

</>

(iii) Programming

Write stable simulation code that handles boundary conditions and renders a video.
Measured by compile / run / video-save rates and CodeBLEU.

⟳

End-to-end

Does the regenerated video show the same physics as the input?
Measured by human ratings of physical plausibility and physics equivalence.

This is not retrieval or memorization: every sample uses randomly sampled parameters drawn from physically plausible ranges, so values must be inferred from visual dynamics. Models get the engine name and an experiment-agnostic code template, so the test measures physics understanding rather than familiarity with an engine.

Physical-law correctness (are the stated laws true, regardless of the video?) versus equivalence (do they match the laws in the video?). The gap between the two bars is the model's grounding gap.

Execute each model's code and compare the regenerated video with the input. Closed models are consistently stronger, and Gemini 2.5 Pro leads every column. Patch-wise metrics are far more lenient than humans.

Of the 2,430 inputs, how often each model's generated code compiles and how often it also runs and saves a video. Models are ranked by the share that produces a video. The gap between the two shades is code that compiles but fails at run time.

runs & saves videocompiles only

Open-source models ran with a 32K-token context window (49K for Pixtral) that has to hold the video frames, the prompt and the answer, which accounts for much of their lower success rate.

Naming recall (fraction of GT parameters the model names) versus value accuracy (fraction of matched parameters within ±20%). Bubble area is end-to-end recovery. Claude and GPT-5 Mini name the most parameters; GLM-Thinking is most accurate on the few it names.

Parameters are matched by greedy bipartite matching over Jaccard similarity on alias sets (lexical and Greek-letter abbreviations), after unit normalization. Scene metadata, geometry, color and g are excluded.

Click a column header to sort. Darker cells are better within a column. Param. Acc. = end-to-end parameter recovery within ±20%.

Frontier modelsTen hard samples, fourteen models

We hand-picked 10 hard samples (5 SciPy and 5 Kubric) and evaluated four newer frontier models, Claude Fable 5.1, GPT-6 Astra, Grok 4.6 and Kimi K3, alongside the ten benchmark models. The frontier models clearly surpass the earlier models. Claude Fable 5.1 shows the strongest physics understanding, and Kimi K3 and GPT-6 Astra produce the most visually similar videos. Even so, parameter recovery stays below 20%.

Law equivalence (overall, 1–5)

Parameter recovery (±20%)

Video similarity (DINOv2 / VideoCLIP)

Side-by-side: input video vs. regenerated simulations

Pick a sample. Every tile plays the video rendered by that model's generated code. Click a tile to see the model's inferred law, its parameter estimates against ground truth, and its code.

Ground truth Input video
Frontier
Closed-source
Open-source

Dataset explorer162 phenomena, one sample each

Filter by physical domain or engine. Hover to play, and click to open the full sample: video, structured chain-of-thought, sampled parameters and the simulation code that rendered it.

162experiments

CitationBibTeX

@misc{kundu2026physimcode,
  title  = {PhySimCode: A Benchmark and Evaluation Method for
            Physics Video to Code Generation},
  author = {Kundu, Souraja and Gupta, Aditya and Julin, Joel and
            Zhao, Yizhou and Xie, Liuyue and Jeni, Laszlo A.},
  year   = {2026},
  note   = {arXiv preprint (coming soon)}
}