Pointing a pair of Ray-Ban Meta glasses at a shelf and asking, out loud, which product has less sugar produces a spoken answer in seconds — no phone, no typing, no photo to review first, just a glance and a question. The interesting AI demos of the past two years increasingly look like this: a camera pointed at something physical rather than a text box waiting for a prompt.

Text-based interaction required a person to first translate their situation into language before a system could help, which filtered out an enormous amount of ambient, physical-world work that never got described in words at all. Multimodal models remove that translation step — but removing it also removes a natural checkpoint where a person previously clarified their own intent by articulating it, so a system that acts on what it sees rather than what it is told risks acting on an ambiguous scene the way a person might misread a glance: confidently and wrong.

Field work — maintenance, retail, logistics, healthcare — stands to absorb multimodal AI faster than desk work, precisely because those jobs generate the physical, unstructured input that text-first tools were never built to ingest. Be My Eyes built its GPT-4-powered assistant, Be My AI, specifically for this kind of hands-free, situational description, aimed at blind and low-vision users who could not type a query in the first place.

Existing product interfaces built around explicit input and output — forms, chat boxes, dashboards — assume a moment where the user actively hands the system a request. Ray-Ban Meta's glasses break that assumption by design, which is why the visibility of their recording indicator light has already drawn scrutiny from privacy researchers who point out how easy it is to miss from across a room.

The products that succeed will likely be the ones that make the boundary of observation explicit and controllable — when the camera is watching, what it is allowed to act on — because ambient systems with unclear scope generate a trust deficit no capability improvement can offset. The wedge was language because language was already digitized. The territory was always the physical world, and a pair of glasses with a camera in the frame is the most literal attempt yet to meet it there.