Comment by KaiserPro
2 days ago
> Colors, amount of daylight(/nightlight), weather/precipitation/heat haze, flowers and foliage, traffic patterns, how people are dressed, other human features (e.g. signage and/or decorations for Easter/Halloween/Christmas/other events/etc.)
I mean, in theory it could. But in practice it'll just output lat, lon and a quaternion. Its going to be hard enough to get the model to behave well enough to localize reliably, let alone do all the other things.
The dataset, yes, that'll contain all those things. but the model won't.
You don't know for sure the model won't contain non-location data, like I noted the additional blurb vaguely said: "And, as noted, beyond gaming LGMs will have widespread applications, including spatial planning and design, logistics, audience engagement, and remote collaboration."
> will have widespread applications
There are a lots of "coulds" "ifs" and "shoulds". But how do you tokenise all those extra bits? For it to function as a decent location system, it has to be "invariant" to weather/light conditions. Otherwise you'll just fall back to GPS.
At it's heart, its a photo -> camera pose (location) converter. The bigger issue is how do you stop it hallucinating the wrong location when it has high uncertainty. That's before you get into scaling issues so that a model can cope with bigger than room scale pointclouds.
the first "public" VPS was released a while ago, yet six years later we still don't see widespread adoption of visual based location, even though its much much more accurate in an urban environment.