In the last week, several reports detailed Apple’s plan to include cameras in the next version of their AirPods. The product remains in development and the material is incomplete, but a few things have been reported. The cameras have both active and passive modes. The active mode includes users requesting information about what they see via Apple’s Visual Intelligence. However, they do not appear to create files for the Photos app. The passive mode is not a continuously-on system, or at least not anymore so than AirPods already are, but it does activate on its own in response to certain conditions.

The information originally came through a MacRumors forum.

The forum member also has more details on “passive mode,” which is described as “lower-resolution and supports environmental awareness.” macOS 26 outlines five possible conditions for a passive capture:

Ambient speech

Audio scene change

Posture change

Head rotation

“Circular region, where the user has moved out of a defined spatial radius.”

The AirPods already sense for and respond to sound and head/body movement. In relation to Siri’s operation, this visual channel combines with other sensory data, including other peripherals. There are new Apple Watches, Pendants, and Glasses reportedly coming with these cameras as well.

To be clear, Apple plainly states that they do not employ user data for things like training new AI models, among others. And, in some respects, they are selling themselves as purveyors of “personal AI,” so privacy is elemental to that. On the other hand, I will say that the kind of data these peripherals would produce is a kind of multimodal data that AI companies are seeking. For example, the NY Times reported on 8/20 about the growing field of data annotation. Workers wear smartphones on their foreheads that record what they are doing with their hands or perhaps while using a pair of robot hands. The point is to acquire more structured visual information.

However, as Apple isn’t doing that with these technologies, what value does“passive capture” provide? On its face, the phone gains more information about the user’s environment, but what does it use that information for? One place to begin is with the iPhone’s current use of ambient information. Location Services uses the phone’s map tracking data to make recommendations about traffic and directions and destinations it predicts you will wish to reach at a certain time. (E.g., You’re usually at work by 9am, you should leave now to arrive on time. Traffic is light.) Attention Awareness checks to see if the user is looking at the phone. The AirPods heart rate monitoring tool shares information with health and fitness apps.

Here is an example, a couple out for a walk, stop briefly, and turn to look at a restaurant menu in a window. One says, “It looks good, but it’s closed today,” The AirPods chirp up that there are three similar restaurants that are open and within walking distance. The head and movement action has activated the passive cameras, which have processed the data from the scene. Siri has inferred that the couple are interested in lunch. It accesses data regarding when the user typically eats lunch from his data, as well as where. It anticipates/predicts the next action.

I am reminded here of Thomas Rickert’s concept of ambient engineering, as he discussed in his 2025 RSQ article. He examines how ambient rhetoric can be effectively engineered, especially now in the context of artificial intelligence. Briefly, he investigates hyper-nudging and hyper-relevance as examples of a subattentive rhetoric. As Rickert suggests, we can think of doomscrolling as an example of hyper-nudging: it is made easy in each moment to continue. Hyper-relevance targets the user as well. Ambient engineering thus becomes useful in exploring how a social media platform moves user populations in affective and political directions.

These devices’ “peripheral vision” also extends posthuman computer vision, which I discuss in a recent journal article. In that article, the focus was on the smartphone screen interfaces as visual and haptic media. However, where that article addressed the expanded, shared capacities of computer vision, in this emerging media ecology of “continuous inference” there is a further shift into how inference becomes actionable, perhaps as a hyper-relevant nudge, to combine Rickert’s terms.

One important point of clarification is to recognize the ways in which inference is continuous and the ways in which it is not. Operationally, if inference is triggered by human movement, as with the AirPods, then it is not continuous. Though there is some passive sensor awake to the call for inference, inference itself is not occurring there. There is continuous inference to the extent that computers in data centers continuously calculate inferences, though not all of them all of the time. Each inference itself requires some provisional end point in order to be output, though the output may be taken up again for further inference.

Another interesting and still unclear aspect of these peripherals is how the data they collect will persist. To continue with my couple seeking lunch. Their location at noon in front of this restaurant has diminishing informational value as time moves on. Other updates replace it. Perhaps it remembers that the couple likes Thai food. Or it is another data point confirming they are often in this part of town at this time. Or that the woman the user is with isn’t his wife. Or whatever. I’m sure it is all completely innocent. They have nothing to hide.

While the technical specifics remain incomplete, these peripherals operate as part of a speculative continuous-inference environment. Like the cameras, microphones, and other sensors distributed throughout the built environment, they anticipate the availability of cheap compute and inference. They make increasingly minor features of everyday life available for computational judgment, including inferences with very little marginal informational value. How much value can the world find in learning once again of one’s preference for Thai food? And yet these devices serve as a discontinuous, peripheral hinge among the world, the user, and an inferential probability space.

Leave a comment

Trending