VRHow / XR Technology / How AR glasses work

How AR glasses work

Short answer: AR glasses add computer-generated light to your view of the real world. A tiny display sends that light through an optical combiner—often a waveguide—while cameras and motion sensors estimate where your head, eyes and nearby surfaces are. Software then renders each frame from the correct viewpoint and anchors it to the scene.

Last checked October 11, 2026. This is a mechanism-focused explainer, not a hands-on review; VRHow has not independently tested the devices mentioned, and capabilities vary by model, software version and region.

1. The image starts in a light engine

Inside an AR glasses system, a small light engine creates the image. Depending on the design, that can involve microdisplays such as LCoS or micro-LED, or laser illumination. The image is not normally shown on a phone-sized screen in front of your eyes. Instead, the engine injects light into the glasses’ optical system.

A waveguide is a transparent optical element that guides and redirects that light so it exits toward the eye while ordinary light from the room continues through. Magic Leap describes its waveguides as transparent displays that control light to overlay digital content on the physical world. Other AR glasses use different combiners, including prisms, beam splitters or “birdbath” optics. The design choice affects thickness, brightness, field of view, eyebox and efficiency—not merely sharpness. A research review in Light: Science & Applications notes that these display goals trade against one another, and that a fixed virtual focal distance can create a vergence–accommodation conflict.

That explains a common surprise: “transparent” does not mean the overlay can be equally bright everywhere. Sunlight competes with the projected image, and the usable viewing area can be smaller than the lens. If your eye moves outside the system’s eyebox, the image may dim, clip or disappear. Field of view is also the window through which the overlay is visible, not the amount of the world the glasses can sense.

2. Sensors turn head motion into a coordinate system

To keep a label attached to a wall while you walk, the glasses need more than a video camera. An inertial measurement unit (IMU) measures rapid changes in acceleration and rotation; outward-facing cameras look for visual features; and some systems add depth sensing. Sensor fusion combines those imperfect measurements into a continuously updated estimate of the glasses’ position and orientation.

Microsoft’s HoloLens 2 hardware documentation is a useful concrete example, listing four visible-light head-tracking cameras, a time-of-flight depth sensor and an accelerometer, gyroscope and magnetometer. That is one device, not a universal AR-glasses specification. Lightweight display glasses may offload tracking and rendering to a phone or compute puck, while some products marketed as “smart glasses” have no spatial tracking or see-through AR display at all.

Applications receive this result as poses: a position and orientation relative to a reference space. The W3C WebXR specification describes spatial spaces and viewer poses in exactly those terms. In practical language, the renderer asks, “Where are the user’s eyes now, and from what viewpoint should this object be drawn?”

3. Mapping gives digital objects somewhere to belong

Tracking the headset is only half the illusion. To place a virtual measuring arrow on a table—or hide a hologram behind a real sofa—the system needs an estimate of the environment. Depth data and camera observations can be turned into surfaces, planes or a triangle mesh. Software can then raycast from a gaze or hand direction, find the table, and attach an object to that surface.

Microsoft’s spatial-mapping documentation explains the practical dependency: mapped surfaces can support placement and occlusion, but surfaces change as the device gathers new data, and an unscanned area cannot be used reliably. A dark, reflective, textureless or moving scene can therefore reduce confidence. “It stayed on my wall” is not a guarantee that every AR system will stay registered in every room.

4. Rendering closes the loop—many times per second

For each display update, software combines the application’s 3D scene with the latest head pose, eye position (when available), display calibration and environmental information. It renders the left and right views with the appropriate perspective, applies distortion or optical compensation, and sends the result to the light engine. The system must do this with low enough latency that the overlay does not visibly lag behind your head.

Eye tracking is optional, not a defining requirement. When present, it can support gaze input and adjust rendering for the user’s eye position. Microsoft’s HoloLens 2 guidance says calibration is required and can fail because of some glasses, contact lenses, eye conditions, sunlight or occlusion. It also recommends fallback input. Treat eye tracking as a capability with permissions and calibration requirements—not as a magic accuracy guarantee.

What the technology means in daily use

What you wantSystem component that mattersLikely limitation
Readable overlay outdoorsCombiner efficiency and display brightnessSunlight can wash out transparent graphics
Object fixed to a deskPose tracking plus spatial mappingTextureless, reflective or changing surfaces can confuse sensing
Comfortable 3D placementOptics, calibration and depth cuesField of view, eyebox and focal-depth trade-offs remain visible
Hands-free selectionEye, hand, voice or controller inputPermissions, calibration, noise and fallback design matter

A useful checklist before believing a demo

For the terminology boundary, see VR vs AR vs mixed reality. For the broader topic map, return to XR Technology Explained or VRHow’s home page.

Sources

Sources were opened and checked for the claims linked above. Product capabilities and terminology change; verify current model documentation before buying or deploying a device.