Skip to main content
GameDev.net gamedev.net
Research Paper

This is an academic paper or technical research. Key findings may require technical background to fully understand.

Explore Research Radar

PRO Tired of ads? Read GameDev.net ad-free and help keep the community independent with GameDev Pro — $3/month.

arXiv cs.GR
arXiv cs.GR Research
· 1 month ago • Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

Briefing

FoldingAgent is a new agentic framework that infers explicit parametric origami procedures directly from demonstration videos. Instead of predicting a static crease pattern, it watches a folding sequence, reasons about the geometry, and outputs an executable folding plan that can be checked against physical constraints.

The system combines a pre-trained vision-language model with specialized tools for simulating geometric transitions, verifying plausibility, retrieving and comparing visual content, and scoring its own predictions. That tool use matters: the agent can re-plan as it goes, which helps reduce the compounding errors that usually show up in multi-step visual reasoning tasks.

The work also introduces PurelandFold, a curated benchmark of Pureland origami videos with ground-truth geometry and action labels. The benchmark gives the approach a concrete testbed for comparing unstructured demonstrations against structured, parametric plans. For developers, the broader takeaway is that video-to-procedure extraction is getting closer to something you can actually build around, not just classify.

While this is an origami-focused system, the underlying pattern is relevant to game tech: sequential inference, tool-assisted validation, and physically plausible reconstruction are all ideas that map cleanly onto animation tooling, procedural content generation, and authoring assistants. If this class of system matures, it could reduce the gap between human-made demonstrations and machine-readable pipelines.

“transform unstructured visual demonstrations into executable, physically plausible folding procedures”

— Maya Moriya et al. · Core claim of the system
Original source
Read on arXiv cs.GR
At a glance
what
FoldingAgent infers parametric origami folding programs from demonstration videos
who
Maya Moriya, Sigal Raab, Yael Vinker, and Tali Dekel
when
Submitted 31 Aug 2026; arXiv:2609.00377
impact
Shows a tool-assisted, sequential approach to turning video into executable procedures
Signal Neutral

Promising research, but still early-stage

Discuss

Follow ai updates

See relevant stories in your personalized news feed.

Sign in to follow

Continue on GameDev.net

Useful next steps related to this story.

Game development news without the noise

One useful weekly briefing. No daily flood.

Sending your confirmation email…

Discussion

Loading comments...