All research

Ways of Seeing, Ways of Naming

Essay7 min read
  • computer-vision
  • vlm
  • governance
  • state-capacity
  • artificial-intelligence
  • ethics

John Berger, Wittgenstein, and the vision-language state.

Axonometric cutaway diagram showing municipal cameras and sensors capturing urban geometry on the left, transforming into semantic text nodes, formula weights, and legible decision trees on the right.
Multimodal perception converts physical observation into structured reasoning: from pixels and bounding boxes to explicit policy weights and contestable administrative decisions.

John Berger begins Ways of Seeing with an extraordinarily simple observation: seeing comes before words.

Before a child can describe the world, it can see it. But Berger’s larger argument is that seeing is not the same as understanding. We do not encounter images neutrally. What we notice, what we call important, and what meaning we assign to what we see are shaped by culture, history, ownership, and power.

An image is never the thing itself. A painting, photograph, or video is a specific way of seeing something. Someone chose the frame, moment, perspective, and subject. Even supposedly objective images contain choices.

Context continuously alters meaning. Place a Renaissance painting in a church and it performs one role. Reproduce it in a textbook, an advertisement, an Instagram post, or a television documentary, and it becomes something else entirely. Mechanical reproduction liberated images from their original locations, but it also rendered their meanings malleable.

Looking is bound up with power. Berger’s analysis of the nude demonstrates how European oil painting systematically constructed women as objects to be observed: "Men act and women appear." The person doing the looking and the person being looked at occupy radically different positions of authority.

Possession shapes Western visual tradition. Oil painting excelled at representing things that could be owned—land, clothes, livestock, food, architecture, and precious objects. The visual tradition developed alongside wealth and property relations.

Modern advertising inherits this exact logic. Publicity does not primarily sell an object; it sells an imagined future version of the viewer. The image creates a gap between who you are and who you might become if you possessed what is depicted.

Half a century later, these insights are unexpectedly crucial for artificial intelligence and the state.

Governments have spent decades building the ability to see. CCTV cameras watch intersections and public spaces. Satellites repeatedly photograph cities. Drones survey roads, construction sites, and infrastructure. Smartphones produce geotagged photographs and videos. Sensors measure air, water, traffic, and weather. What began with occasional surveys and photogrammetry has progressed to orthomosaics, point clouds, and increasingly rapid 3D reconstruction.

The problem of capturing reality has not disappeared. Scale, frequency, resolution, and cost still matter enormously. But technologically, something fundamental has changed: it is increasingly possible to create a continuously refreshed representation of the physical world.

Computer vision has given machines the ability to parse some of what they see. A model can identify a road, vehicle, building, tree, pothole, or patch of standing water. Vision-language models go one step further: they translate visual scenes into language.

A machine can look at an image and say, in effect: this road is damaged; there is water accumulating here; construction has occurred on this parcel; this tree canopy has disappeared.

We are beginning to give words to everything the state can see.

Modern artificial intelligence reverses Berger’s opening premise. Berger starts with the proposition that seeing precedes words. Vision-language models move in the opposite direction: turning seeing into words. A camera captures a scene; computer vision identifies objects and spatial relationships; a VLM converts what is visible into linguistic concepts that a computational system can reason over.

This raises a fundamental question: whose "way of seeing" gets encoded in the model? A CCTV frame may be physically objective in the narrow sense that photons hit a sensor, but deciding that a cluster of pixels represents encroachment, congestion, unsafe infrastructure, a pothole, a crowd, or suspicious behavior is already an interpretation. The model isn't merely observing reality—it is imposing a vocabulary on reality.

This collapses the insights of Berger and Wittgenstein into a single computational engine. Berger asks how our ways of seeing determine what the world appears to contain; Wittgenstein asks how the limits of language determine what can meaningfully be said about that world. Vision-language models merge both problems into one system.

"The limits of my language mean the limits of my world." — Ludwig Wittgenstein

For computational systems, seeing and naming are now the same process. Once machines can both see the world and give language to what they see, the next question is no longer whether they can observe reality.

It is whether they can help us decide what to do about it.

From Seeing to Deciding

Consider something relatively mundane: potholes.

Suppose a government has cameras mounted on municipal vehicles, smartphones carried by field workers, and video collected from vehicles moving through the city. Computer vision processes these feeds and identifies 40,000 road defects.

This is a considerable improvement over a system dependent on complaints or periodic manual inspections. For the first time, a city might have something approaching a systematic picture of road conditions.

But immediately a harder question appears: which pothole should be repaired first?

Computer vision cannot answer this merely by becoming more accurate. The system can measure defect dimensions, estimate severity, and identify road category. It can combine this with traffic volumes, accident histories, public complaints, nearby schools and hospitals, population served, and repair costs.

It can know an extraordinary amount. But moving from what exists to what matters requires something fundamentally different.

There is no objective answer hidden inside the pixels.

Imagine two roads: one is a badly damaged village road used by 800 people every day; the other is a moderately damaged urban arterial used by 80,000. Prioritising the arterial maximises the number of people affected. Prioritising the village road addresses a more severe deprivation. Weighting accident probability produces another answer. Weighting economic productivity produces another. Giving historically underserved neighbourhoods additional priority changes the ranking again.

At this point, the problem has stopped being a computer-vision problem. It has become a problem of values.

There Is No View From Nowhere

A camera appears neutral because it records what sits in front of it. But even before an algorithm enters the picture, choices have already been made:

  • Where was the camera installed?
  • Which neighbourhoods are surveyed, how often, and at what resolution?
  • What happens at night?
  • What data is retained?
  • Which parts of the city have enough digital infrastructure to become visible in the first place?

Reality capture can therefore be accurate without being complete.

Pixels do not inherently contain "encroachment"—they contain shapes, colours, and spatial relationships. Humans define a category called encroachment and establish the conditions under which something belongs to it. The same is true for unsafe building, illegal dumping, congestion, road defect, crowd, flooding, or high-risk infrastructure.

A vision model does not escape Berger’s way of seeing. It industrialises it. We teach the machine what distinctions are worth making—and then ask another machine to determine what those distinctions should cause us to do.

Neutrality Is the Wrong Goal

There is a strong temptation in government technology to describe algorithmic decision-making as neutral.

Human decision-making can be inconsistent. Officers interpret rules differently, political pressure alters priorities, files move through personal connections, and loud complaints drown out larger but less visible problems. An algorithm appears to offer an escape: give every problem a score, rank them, and allocate resources from the top down.

The neutrality is partly an illusion.

Suppose road-repair priority is calculated by formula:

VariableWeightPolicy Intent
Safety Risk40%Minimise accident and injury probability
Population Affected25%Maximise benefit per user
Defect Severity20%Prevent catastrophic structural failure
Economic Importance10%Support commercial throughput
Repair Cost5%Optimise fiscal efficiency

The computer can apply this formula with total consistency. But the formula itself is not neutral.

Why is safety worth 40% rather than 30%? Why does economic importance receive 10%? Should cost even reduce the priority of repairing something dangerous? Should poorer neighbourhoods receive an equity adjustment? Should complaints matter if wealthy neighbourhoods are more likely to submit them?

These questions cannot be resolved by better machine learning. They are questions about what society values.

From Neutrality to Legibility

The most important property of an AI decision-making system may not be neutrality, but legibility.

Imagine that instead of simply producing Road A — Priority 1, the system explains:

Road A ranked above Road B because its estimated accident risk is 2.4 times higher, it serves 31,000 daily users, its defect severity exceeds the intervention threshold, and current municipal policy assigns 40% of the priority score to safety.

Now something important has happened: the system has exposed its chain of reasoning.

A citizen, engineer, auditor, or elected representative can ask whether the underlying image was correct, whether the model correctly classified the defect, whether the traffic estimate holds, or whether safety should carry 40% of the score.

The system no longer asks for trust based on algorithmic objectivity. It exposes what it saw, how it interpreted the scene, what evidence it attached, what rules it applied, and what values those rules embedded.

That chain can be inspected—and, crucially, it can be contested.

The Decision Becomes an Object

Today, many government decisions are difficult to interrogate because their reasoning is distributed across siloes.

Observations sit in inspection reports, population data lives in another department, complaints sit in a CRM, budget limits live in a spreadsheet, and political instructions arrive by phone. The final decision emerges from this mixture, but the path that produced it is rarely reconstructed.

AI creates the possibility of making the decision itself a structured, observable object:

Observation → Interpretation → Evidence → Policy → Weighting → Recommendation → Human Action

Every step can be timestamped. Every source can be referenced. Every model version can be recorded. Every policy weight can be visible. Every human override can require a documented reason.

This does not eliminate discretion—it makes discretion observable.

A Different Kind of Accountability

Once the reasoning chain becomes visible, administrative disagreements become precise.

Suppose residents argue that their neighbourhood systematically receives lower infrastructure priority. Without a legible system, this becomes an argument about bad faith or neglect: Is the government ignoring us?

With a legible system, the question becomes empirical:

If complaint volume carries 20% of the prioritisation score, and wealthier neighbourhoods generate three times as many digital complaints, the system produces an unequal outcome without any explicit discriminatory rule.

The disagreement can now move upstream.

Should complaint volume carry less weight? Should detected physical condition carry more? Should low-reporting areas receive an equity offset?

The political argument does not disappear—it becomes visible in the architecture of the system.

Berger's Machine

Berger’s insight was that images do not simply show reality; they carry a way of seeing.

AI does not eliminate this problem—it makes the way of seeing executable:

  1. The camera decides what enters the field of vision.
  2. The vision model decides what can be recognised.
  3. The language model decides how observations are described.
  4. The data system determines what administrative facts connect to them.
  5. The decision model determines what matters.
  6. The policy determines how much it matters.
  7. The institution determines whether anything happens.

There is no neutral answer to how a state should distribute scarcity. But there can be a system in which the state's way of seeing is explicit.

The State's New Interface With Reality

For decades, the state’s relationship with the physical world has been episodic. Surveys, complaints, inspections, and censuses happen periodically. The physical world changes much faster than the administrative databases built to describe it.

Continuous reality capture—satellites, drones, CCTV, mobile video—alters that cadence.

Reality → Seeing → Naming → Understanding → Prioritising → Acting

The first four stages are rapidly becoming technical capabilities. The last two never entirely will be—and shouldn't be.

The purpose of AI in government is not to hide political judgment inside opaque models, but to expose the relationship between evidence, judgment, and action.

Berger taught us that there is no seeing without a way of seeing. The opportunity with AI is not to build a machine with no point of view, but to build institutions where the state's way of seeing—and the path from what it sees to what it does—can finally be seen.