The most commercially successful smart glasses of the last few years were not the ones with the most technology in them. They were the ones people were willing to wear. That distinction is worth examining closely, because it says more about the future of the category than any specification comparison.
More than a million units were sold in the first year of the current generation. The previous generation, launched two years earlier with a nearly identical product definition, sold about 300,000 units over its entire two-year life. The newer pair matched that figure in a single quarter. The difference between the two products, at the level of description, is one additional gram of weight and a better camera.
Why the frame came first
The generation that failed was not a bad product; it was built on the same Ray-Ban frames with the same camera and audio functions. It failed because the two functions were not yet good enough for anyone to change their habits for, and habit is the only thing that keeps a wearable on a face past the first week.
The pair that succeeded started from a different premise. Its claim to adoption is that it is first of all a serviceable pair of sunglasses. That is a low bar in absolute terms and a very high one for a technology product. It means the frame has to look right to people who do not care about technology, and it has to be comfortable enough that the wearer stops noticing it.
On the first count, the design comes from an eyewear company with a decades-long record, and the smart variant differs from the conventional frame only at the front corners. The available shapes - a squared frame, a rounded frame and a cat-eye, each in several colours and with a choice of tinted, photochromic, polarised or clear lenses - meant buyers could choose an appearance they would have chosen anyway.
On the second count, the official weight is 48 g. That is at the upper end of what most people tolerate in a frame, and heavier than ordinary spectacles. In practice the pair slides more often than a conventional one, and some wearers never stop noticing it. The relevant point is that the discomfort does not cross the threshold at which the device gets left at home: it settles into a register of "a slightly heavy pair of glasses", which is a different category of complaint from "unwearable".
The temples keep the tapered curve of the original frame, which required a custom speaker geometry rather than a standard rectangular or oval driver. That detail is the clearest illustration of the priority this product puts on appearance over internal convenience.
What makes the pair worth carrying
Three functions justify the weight, and only one of them is artificial intelligence.
Open-ear audio was the first. Sound quality is comparable to a decent set of open-ear headphones, which is to say good enough for music in a quiet room and for podcasts anywhere. Because the ears are not occluded, the wearer remains aware of traffic and of other people, which is the reason the format is appropriate for cycling and for walking in a city in a way that in-ear headphones are not.
The camera is the second and the more transformative. The scenes it captures are the ones a phone cannot: a view seen while driving, an animal that will not wait, a child's expression that lasts half a second. The current camera is a 12 MP unit capable of clips up to three minutes with noticeable stabilisation, up from a 5 MP sensor in the previous generation that was, in practice, barely usable. Recording is triggered by a button on the frame or by voice.
Audio capture uses a five-microphone array, which is more than most dedicated recorders carry, and the spatial result is part of why the footage is worth keeping. The third function follows from the first two: once a device is always on the face, it can be used for the recording that would otherwise not happen at all.
The behaviour that emerges from those three functions is the finding that matters. Owners who were surveyed before this purchase did not describe the glasses as a gadget they used occasionally; they described a settled habit. A gym user who listened to music through the frame explained it simply: there is always a chance something worth photographing will appear. That reasoning recurs across every owner, and it is the mechanism by which a technology product becomes a daily object.
Where that camera sits against three current competitors, and what each of them gives up to get there, is set out in AI glasses for video.
What the assistant adds, and where it stops
Voice control existed in the previous generation in a very limited form: basic volume and call commands through a connected phone. The change that turned the product into an AI device was the addition of a multimodal model, which made it possible to ask a question about what the camera sees rather than only about text.
Latency is the most impressive part. A short processing tone plays after each question, and in practice answers arrive within roughly one tone, whether the question is visual or textual. The likely explanation is on-device preprocessing of the image into tokens before the request is sent, on hardware that was designed with this workload in mind.
Capability is the less impressive part. Answers are length-limited and stop abruptly, object recognition is inconsistent, and counting objects in a scene remains unreliable. Questions of the "where can I find water in this room" variety are answerable but essentially artificial; the genuine use cases are narrower than the demonstrations suggest. Live translation was added later and works, with a noticeable delay.
The argument for the assistant is not that it is capable. It is that the delivery mechanism is different. Speaking to a phone requires taking it out, and speaking to a home speaker requires being at home. A pair of glasses that is already on the face and already closest to the mouth makes the same request cost less, and a request that costs less gets made more often. Users who had habitually used a phone assistant for music playback describe switching within a week, and the request patterns that followed were ones nobody had planned. They included a poem on a boring journey, an opinion on an outfit in a mirror, and inspiration for a meal.
Some of what was a software limitation in this generation has since been revised, and the product line is documented under Ray-Ban Meta.
A third path for the category
The industry has spent more than a decade assuming that glasses would eventually replace the phone as the primary interface, and the history of that effort explains why this product is interesting.
The first attempt put a display in the frame and tried to reproduce phone functions through it. The prototype was compelling and the product was not: battery life and heat made it impractical, and the functions it actually delivered were a small subset of what it promised. Two schools followed. One treated glasses as consumer electronics peripherals - cameras, or audio devices - and stopped claiming to replace anything. The other pushed display technology harder, moving more and more computing into head-mounted devices, at the cost of weight and of use cases confined to a room.
Both schools missed the property that made glasses promising in the first place: they are worn continuously and they sit closest to the eyes, ears and mouth. A headset with excellent optics is not worn on a bus. A pair of glasses with a camera and a model behind it is.
The technical prerequisites now exist. Chip design has brought the power draw of continuous capture within reach of a small battery, and camera quality has improved enough that the footage is worth reviewing. A language model makes voice-only interaction viable in a way that it was not before. That last point is the substantive change: voice interaction was previously too weak to carry a platform, and it no longer is.
The near-term promise is therefore not a replacement for the phone. It is a device that handles a growing set of interactions - messages, questions, capture, payments, navigation - through speech and vision, without being taken out of a pocket. The interface for those services largely exists; what is missing is the platform work to connect them. That is a commercial problem rather than a technical one, and it is the reason this category is worth watching again.
The path that depends on a display rather than on capture and audio is surveyed in Display AR glasses compared.