부산대학교 그래픽스 및 기하 처리 연구실

PNU Graphics & Geometric Processing Lab
School of Computer Science and Engineering, Pusan National University
* * *
Research / 04

How VLMs read 3D shape

Study the evidence in a rendered image by varying how surface faces are revealed and how recognition is measured.

01 / A model receives images, not a mesh

A rendered 3D recognition task includes a graphics pipeline before the model ever sees its input. Camera position, selected faces, projected silhouette, shading and answer format can all change the available evidence.

Progressive surface disclosure makes this dependency explicit. A sequence reveals increasing portions of a mesh and asks when a model first identifies the object correctly. That first-correct budget measures performance under the chosen stimulus and scoring protocol.

Revealing a few faces on distant parts of an object may expose its overall arrangement before any one part is complete. Revealing the same area around one location may instead show a locally coherent piece whose identity remains ambiguous. This difference motivates varying the disclosure ordering while matching the nominal amount of surface.

A threshold from this procedure is attached to an entire sequence of trials. It depends on which partial views came first, what the model was asked, and how the response was judged. It should therefore be interpreted together with those choices, rather than as an intrinsic percentage of the object that a model needs to understand.

* * *

02 / Match area, vary its distribution

Random and connected face disclosure at matched area budgets

Computed on the same synthetic closed vase as the original figure. Six views are front, back, left, right, front-left-top and front-right-top (read across each row). The single view maximizes full-mesh projected area once, then stays fixed. No hidden faces or complete-model outline are drawn. Silhouette mode flattens every selected-surface pixel to gray. Coverage matching finds each ordering’s shortest prefix reaching Random’s image coverage, measured on fixed 128 × 128 masks per view. This browser demonstration uses illustrative cameras and raster settings, not the paper’s original renderer or object set. Percentages are rounded to integers; 100% is an added full-shape reference.

The budget is defined by triangle surface area, not triangle count or the number of image pixels. Given an ordering of faces, the visible set is the shortest prefix whose accumulated area reaches the target fraction of the total mesh area.

OrderingSelection rule
RandomA fixed random permutation of faces.
Area-firstFaces in descending triangle area.
PCA sweepFace centroids ordered along the first principal axis.
Connected growthExpand over face adjacency, starting from a largest face and prioritizing area.
Random connected growthUse fixed random priorities for the starting face and frontier.

A graph-connected set need not look compact in one view: connections can be thin or occluded. Likewise, equal 3D surface area does not guarantee equal image-space coverage.

The prefix construction makes each sequence progressive: increasing the budget retains the previously selected faces and adds more. Because a whole triangle is the smallest selection unit, the realized area can slightly exceed the target. The experiment checks realized areas so that an ordering effect is not confused with unequal surface exposure.

PCA sweep and connected growth impose different forms of spatial structure. A sweep follows a global axis and can reveal disconnected parts at similar coordinates. Connected growth follows adjacency on the mesh, while its area-based frontier can still form thin paths or irregular regions. Random connected growth changes the seed and frontier priority to test whether the large-face preference explains the observed behavior.

* * *

03 / Specify the complete trial

Selection, rendering and scoring as three parts of the experimental protocol
Changing the ordering alone is a controlled comparison only when the other trial conditions are held fixed.

The study evaluates 100 everyday meshes across five local VLMs. Its main rendering condition presents six canonical views in a 2 × 3 grid. A separate single-view condition uses one fixed view chosen for each object.

The answer formats are also distinct: object-choice, category-choice and free-response naming. These should not be treated as interchangeable recognition tasks, because the options and scoring rules change what counts as success.

The six-view grid makes several projections available in one query. Keeping these views fixed across a sequence prevents the camera from changing the evidence at the same time as the face set. The single-view control addresses a different question: whether the pattern persists when the model is given only one of those projections.

Object-choice offers a known set of candidate identities, category-choice asks for a broader class, and free-response requires the model to produce a name. Free-response results also depend on whether a fixed alias list accepts equivalent names. Reporting the task and scoring rule is therefore necessary to interpret a correct answer.

* * *

04 / Interpret the threshold with its controls

Random face disclosure yields lower first-correct thresholds than sweep-like or connected disclosure in the tested conditions, with the same direction across all five evaluated models. Changing the answer format also produces substantial differences.

A forced-choice model can occasionally guess correctly. The study therefore pairs censored thresholds with label-shuffle chance comparisons. Silhouette-only and coverage-matched controls further test whether the ordering effect is attributable to interior shading or simply to the amount of projected evidence.

Repeated forced-choice trials create repeated opportunities for a chance hit. A first correct response alone can consequently overstate the evidence for recognition. The label-shuffle comparison supplies a chance reference for the particular task, while censoring specifies how sequences that never succeed contribute to the reported threshold summaries.

The rendering controls ask which visual explanation remains plausible. Silhouette-only trials remove interior shading cues; coverage-matched comparisons address the quantity of projected evidence. The reported ordering effect survives these controls, supporting an interpretation involving where evidence is distributed in the image rather than only how much surface or shading is visible.

A protocol-dependent measurementThe first successful answer is an operational threshold, not proof of complete 3D understanding. Comparisons should report the disclosure rule, budgets, views, answer format, scoring and handling of unsuccessful sequences. The controls support an interpretation involving distributed projected shape within these experiments.
* * *

Related paper

* * *