4.1 Capturing Data for Structure from Motion

4.1.1 Introduction to Structure from Motion

Structure from Motion (SfM) is a photogrammetric technique that combines the accuracy of photogrammetry (i.e. art, science and technology of obtaining reliable information about physical objects and the environment through the process of recording, measuring and interpreting photographic images/American Society of Photogrammetry and Remote Sensing) with the details and automation of object and shape recognition and modelling that computer vision provides (Szeliski 2010). The two methods also start from two opposing premises; on the one hand, photogrammetry starts from the camera properties and then moves to the identifcation of common points in images, whereas on the other hand, computer vision starts from analysing the images and then moves to the reconstruction of geometry (Aicardi et al. 2018, p. 258). However, despite the fact that photogrammetry and computer vision were developed for different purposes, i.e. for metric accuracy in mapping and for automatic extraction of data from images that are machine readable, respectively, their aims started to converge, especially in the late 90s/ early 2000s when image-based modelling and computational photography were rapidly developing.

Structure from Motion combines both photogrammetry and computer vision methods of analysis, aiming to reconstruct both the position of the cameras as well as the three-dimensional geometry of the captured scene/object. By analysing, therefore, the sequential change of a camera position relative to the subject in an image dataset, SfM can determine the 3D structure of the photographed subject. To determine the camera position in each photograph, SfM algorithms look at individual pixels trying to identify the same pixel in more than one photograph. 
LaserScanningVSPhotogrammetry

Laser Scanning vs. Photogrammetry (click to enlarge the figure)


Recent advances in cameras, computer processors and photogrammetry/SfM software, have turned photogrammetry into a powerful technique which produces dense and accurate models, similar to those produced by non-contact 3D digitisers (laser scanners); devices that capture three-dimensional information by using laser or light projection techniques to collect points at a high rate and produce results in real time. In the image on the left, you see a comparison between laser scanner and photogrammetry/SfM applications. When it comes to the cost of the two, laser scanners cost significantly higher than a camera required for photogrammetry. Depending on the resolution that laser scanners can capture, their cost might range from a few thousand euros to over €50.000. At the same time, the processing software is often proprietary (which comes with a higher cost) and their learning curve quite steep. Quite on the contrary, photogrammetry only requires regular camera equipment and software that is either free or with relatively low cost. Although the most prominent photogrammetric software is also proprietary and increasing number of free/open source solutions can compete with the results that commercial software produces. Although the results of the two are nowadays quite comparable (centimeter or sub-centimeter accuracy) laser scanners are quite dependent on their distance from the subject, its size and material, as well as texture. On the other hand, photogrammetry is relatively independent when it comes to the distance between the camera and the object as well as their dimension. Anything that can be photographed, can work well in photogrammetry applications. Both methods have some weak points when it comes to certain materials. For example, photogrammetry cannot work in environments without light or for objects/surfaces that are shiny and highly reflective. On the other hand, laser scanners also cannot cope with reflective or black surfaces; and many laser scanners are not good in capturing textures. This is mainly because scanners are devices originally developed for engineering, in which accurate geometry (and not texture) is the most important factor. Depending on the type and size of the subjects (e.g. objects with rough edges), data acquisition via laser scanners can be very time consuming (especially with hand-held scanners) contrary to photogrammetry that would typically take a few minutes for the trained user. Similarly, processing laser scanned data (points clouds) is typically more time consuming than processing a photogrammetric dataset. Both are also dependent on the computational power of the device used for processing. Laser scanners are powerful devices that can achieve millimeter or even sub-millimeter accuracy. 

Until a few years ago, there were no alternatives to laser scanning and photogrammetric/computer vision technologies were not advanced enough to provide comparable results. Both methods have strengths and weaknesses and therefore any decisions regarding the employment of one technology over the other should take into account the needs of the individual/institution and the peculiarities and future uses of the recorded subjects. Choices driven by technological fetishism and/or superficial knowledge of the strengths and exigencies of the technologies will most likely lead to poor results and implementations.                 

Since Structure from Motion is a universal method that can be applied to any kind of object there is a broad range of potential application areas beyond cultural heritage, including engineering surveying and civil engineering, industrial applications, medicine (Luhmann et al. 2019) and forensics (Chapman and Colwill 2019). The slideshow below presents some characteristic SfM applications, mostly focusing on cultural heritage and related fields. Scroll with your cursor in the window to see the full content.





References



Further Readings