A Benchmark and Evaluation Method for
Physics Video → Code Generation
✉ Corresponding author & project lead: sourajak@cs.cmu.edu
Can a multimodal LLM watch a 2D or 3D physics simulation, work out the involved physics laws and parameter values, and write code from scratch that regenerates it?
PhySimCode is a benchmark of 160.6K procedurally generated physics videos covering 162 phenomena, rendered by two different engines (SciPy in 2D, PyBullet in 3D). Every video comes with the code that made it, its true parameter values and a step-by-step physics explanation, so a model's answer can be checked at every level.
Most models state physical laws correctly (7 of 10 score >4.5/5), but the best (GPT-5 Mini) scores only 2.47/5 at recognizing the law demonstrated in the video.
The best model (Claude Sonnet 4.6) names and estimates within ±20% only 7.7% of all ground-truth parameters.
DINOv2 and VideoCLIP agree with each other (κ = 0.717) but barely with human judgments of physics equivalence (κ ≤ 0.205). Models match appearance and miss dynamics.
162 experiments × 15 uniformly random samples (seed 42). Law scores are averaged over three LLM judges (Gemini-2.5-Flash, Grok-4.1-Fast, GPT-5-Nano). Human scores are the mean of four expert raters over 100 random samples per model. Hover a chart for exact values; click legend entries to toggle them.
Extract visual features and track objects across frames.
Measured by human appearance equivalence and DINOv2 / VideoCLIP similarity.
Bind what it sees to what it knows: identify the governing law and estimate parameters.
Measured by law correctness vs. equivalence and parameter recovery (±20%).
Write stable simulation code that handles boundary conditions and renders a video.
Measured by compile / run / video-save rates and CodeBLEU.
Does the regenerated video show the same physics as the input?
Measured by human ratings of physical plausibility and physics equivalence.
This is not retrieval or memorization: every sample uses randomly sampled parameters drawn from physically plausible ranges, so values must be inferred from visual dynamics. Models get the engine name and an experiment-agnostic code template, so the test measures physics understanding rather than familiarity with an engine.
Physical-law correctness (are the stated laws true, regardless of the video?) versus equivalence (do they match the laws in the video?). The gap between the two bars is the model's grounding gap.
Execute each model's code and compare the regenerated video with the input. Closed models are consistently stronger, and Gemini 2.5 Pro leads every column. Patch-wise metrics are far more lenient than humans.
Of the 2,430 inputs, how often each model's generated code compiles and how often it also runs and saves a video. Models are ranked by the share that produces a video. The gap between the two shades is code that compiles but fails at run time.
Open-source models ran with a 32K-token context window (49K for Pixtral) that has to hold the video frames, the prompt and the answer, which accounts for much of their lower success rate.
Naming recall (fraction of GT parameters the model names) versus value accuracy (fraction of matched parameters within ±20%). Bubble area is end-to-end recovery. Claude and GPT-5 Mini name the most parameters; GLM-Thinking is most accurate on the few it names.
Parameters are matched by greedy bipartite matching over Jaccard similarity on alias sets (lexical and Greek-letter abbreviations), after unit normalization. Scene metadata, geometry, color and g are excluded.
Click a column header to sort. Darker cells are better within a column. Param. Acc. = end-to-end parameter recovery within ±20%.
We hand-picked 10 hard samples (5 SciPy and 5 Kubric) and evaluated four newer frontier models, Claude Fable 5.1, GPT-6 Astra, Grok 4.6 and Kimi K3, alongside the ten benchmark models. The frontier models clearly surpass the earlier models. Claude Fable 5.1 shows the strongest physics understanding, and Kimi K3 and GPT-6 Astra produce the most visually similar videos. Even so, parameter recovery stays below 20%.
Pick a sample. Every tile plays the video rendered by that model's generated code. Click a tile to see the model's inferred law, its parameter estimates against ground truth, and its code.
Filter by physical domain or engine. Hover to play, and click to open the full sample: video, structured chain-of-thought, sampled parameters and the simulation code that rendered it.
@misc{kundu2026physimcode,
title = {PhySimCode: A Benchmark and Evaluation Method for
Physics Video to Code Generation},
author = {Kundu, Souraja and Gupta, Aditya and Julin, Joel and
Zhao, Yizhou and Xie, Liuyue and Jeni, Laszlo A.},
year = {2026},
note = {arXiv preprint (coming soon)}
}