The same footage costs different people different amounts
Being photographed on a street is an inconvenience for most people and a serious risk for some. Where you were, who you were with, and whether that fact reaching the wrong person changes your housing, your employment or your safety are all variables, and they do not distribute evenly.
Scholarship has started to look at specific populations rather than the aggregate. Writing in The Conversation earlier this year, Brynn Colledge examined what facial recognition on smart glasses would mean for sex workers and other people already exposed to targeted harm. Work like that is the useful direction, and it is worth reading in full rather than summarised, because the details are the argument.
Policy calibrated to the median is calibrated wrong
Almost every proposal in this space asks whether a safeguard is proportionate, and answers by imagining a reasonable person mildly bothered by a camera. That person is a construct. The distribution has a long tail, the tail is where the actual damage sits, and a rule tuned to the middle of the curve does nothing for the end of it.
This is not an argument for treating every encounter as maximum risk. It is an argument that the unit of analysis is wrong. Asking what a typical bystander loses produces a permissive answer every time, because the typical case is genuinely mild.
What we would change in how the question is asked
We would evaluate any proposed safeguard against the worst realistic case in the room rather than the median one, the same way accessibility work stopped asking what the average user can perceive.
That reframing has a cost we are not going to hide. It makes safeguards look expensive relative to the harm most people experience, and somebody always points that out. The response is that a public space is shared with everyone in it, including the people for whom a recording is not a minor annoyance, and a rule that serves only the comfortable majority is not a rule about public space at all.