All research

Building Without Blocks

Essay2026 · 0726 min read
  • deep-tech
  • govtech
  • india
  • digital-public-infrastructure
  • ai
A surveying drone hovering over a dam construction site in a dark river valley, illuminated by a warm orange laser light.
Surveying the Polavaram submergence zone: mapping physical reality where digital records stop short.

The standard reading of Indian deep tech is a catch-up story. Funding arrived late, capability arrived later, and the country is now behind in the AI race. That reading is not wrong. It is measuring the wrong thing.

The Dam

The Polavaram project had been in the plans for roughly fifty years before I stood on it.1 By the time I did, it was finally under construction, and a drone was flying over the submergence zone to survey it. I had just moved back to the country, and I was there to help turn what the drone saw into something a planner could use.

The engineering was not the hard part. The hard part was people. Rehabilitation and resettlement had been carried out on this land repeatedly, across decades, against records that described villages which had since moved, villages which had not, and in a few cases villages that appeared to exist in two places at once. A woman on the ground said the thing everyone in every Indian policy room says, that the heart of India lives in her villages. She was not wrong either. It just did not describe the settlement she was standing in, which was not the village in the record, whose compensation had already been paid once, and over which the drone was at that moment generating what would become the fourth mutually incompatible map of the same piece of earth.

Here is the part that took me years to understand. Everything about Polavaram was known. The submergence extent was known. The affected population was known in outline. The delays, the backlog, the disputes, all known. None of that knowledge produced motion at anything close to the speed the knowledge justified.

I did not have the words for it then. The rest of this essay is the words.

Two Indias, One Signal Bar

The strangeness that motivates everything I am about to argue is this. I had returned to a country whose digital public infrastructure was genuinely ahead of most middle-income countries and a few rich ones. Payments worked. Identity worked. Benefit transfer worked. There was 5G in places I did not expect it and internet penetration deep into places I did not expect it either. And a district engineer trying to establish where a road was supposed to go was working from a paper trace of a survey conducted decades before he was born.

Indian technology discourse conflates two things constantly, so let me separate them. Digitisation of records means the register is now a database. Digitisation of reality means the database corresponds to the ground. India did the first at remarkable speed and has barely started the second.

This matters more than it sounds, because digital public infrastructure is a transaction layer. It moves money, verifies identity, and completes exchanges between parties who already agree on what is being exchanged. It does not describe physical reality. And every genuinely hard problem in Indian governance is a physical reality problem: land, water, encroachment, construction, delivery to a specific person at a specific place. The success of the transaction layer made the absence of the physical layer harder to see, not easier, because it created the impression that the digitisation problem had been solved.

Without a trusted spatial substrate, every downstream system is a negotiation rather than a computation. Two agencies with two base maps do not have a technical disagreement, they have a political one, and no amount of software resolves it. This is why so much of what I built worked in demonstration and failed in deployment. The demonstration ran on one version of reality.

Axonometric cutaway showing three stacked layers of spatial data over a river valley.
The spatial substrate stack: bridging paper traces, drone point clouds, and live governance maps.

Ola Is Not Uber, and Blinkit Is Not Anything

Before I can make the argument I want to make, I have to deal with the story everyone already believes about Indian innovation, because a reader who holds it will not accept a rebuttal that has not been fair to it first.

The story goes like this. Indian venture capital spent a decade funding copies. Uber exists, therefore Ola. Amazon exists, therefore Flipkart. The only real delta was operational expertise and knowing Indian geography better than a foreign entrant could. Distribution, not invention. It was a good business and a poor kind of innovation.

A great deal of what got funded fits that story. I am not going to pretend otherwise.

But the same decade produced quick commerce, and quick commerce is not a copy of anything. Ten-minute delivery is unavailable across most of the world at any price. It is absent not because the idea is difficult to have but because the inputs to build it do not exist elsewhere. Dense cities. Dense and available labour. Informal logistics that can be formalised at the edges without being rebuilt from scratch. A customer base that will trade margin for immediacy. Swiggy and Blinkit are genuine innovations. Their innovation happens to sit in labour orchestration and micro-warehousing under density rather than in a novel algorithm, and that does not make it a lesser kind of invention. It makes it a different kind, native to the inputs actually on hand.

Sit with that, because it is the whole essay in one example. India built world-class systems out of the inputs it had, and did not build them out of the inputs it lacked.

The Endowment Argument

Here is the mechanism, stated plainly.

India innovates genuinely, and sometimes world-leadingly, at the layer where its factor endowments are unusual. It innovates derivatively, or not at all, at the layer where its endowments are ordinary or absent.

This sounds obvious and almost nobody applies it consistently, which is why the debate about Indian innovation keeps producing bad answers in both directions, triumphalist and defeatist, sometimes in the same week.

Unusual Indian InputsOrdinary or Absent Indian Inputs
Labour cost and availabilitySilicon
Urban density and informalityPrecision components
Scale and heterogeneity of demandSensors and optics
Linguistic diversity, as a forcing functionHigh-resolution earth observation
Tolerance for operational complexityFrontier research base
Digital public rails, unusually earlyCompute, and patient capital

Run the test across sectors and it holds. Quick commerce, UPI, and Aadhaar sit cleanly in the left column and are world-leading. Foundation models, precision components, and satellite imaging sit in the right column and are not. Drone-based survey, the thing I spent a decade on, sits awkwardly across both, which is exactly why its story is neither triumph nor failure but something more useful than either.

I want to head off the obvious objection, which is that this sounds like an excuse. It is not, and the proof is that the argument predicts where India should be criticised, and the criticism it produces is sharper than the usual one. The usual complaint is that India has not built a frontier AI lab. But compute and research depth are inputs that take a generation and enormous capital to create, and criticising their absence is close to criticising geography. The component tier is different. Sensors, optics, actuators, precision machining, test and qualification infrastructure: these are absent inputs that a decade of policy and capital could have created and did not. That is the failure worth naming. Almost nobody names it, because a component tier produces no announcement, no ribbon, and no single photograph.

From this one mechanism, three consequences follow. They organise everything I saw on the ground.

Consequence One: You Build the Blocks Yourself

A consumer internet company founded in India in 2016 could buy nearly every capability it lacked. There was a supplier tier, an agency layer, a logistics market, a pool of people who each knew one thing well. You assembled.

A deep tech company founded the same year could buy almost none of it. No component base, no fab, no sensor supply chain, no trained operator pool, no domain-fluent integrator layer, no acquirable team, and no capital willing to fund the years it takes to build a block that nobody yet wants.

So we built the blocks ourselves. That is not a strategy, it is a sentence you serve. You set out to sell an analytics layer and discover you must first build capture, then processing, then a base map, then a training pipeline, because none of them can be bought. Four layers built in order to sell the fifth. Time to first revenue stretches past the patience of the capital that funded you, not because execution was slow but because the scope was quadruple what was underwritten. Credibility can only be accumulated, never purchased, because there is no third party the buyer can benchmark you against, no analyst category, no reference implementation. The only proof is years survived.

The selection pressure this creates does not favour the best technologists. It favours the most solvent and the most patient. It produced a founder generation trained in endurance rather than in speed, which was exactly the right training for 2018 and is a live liability in 2026, when the AI cycle rewards the opposite disposition. I will come back to that, because it is the least comfortable thing in this essay and I am inside the question rather than above it.

There is a specific version of this that everyone in the sector lived. You build hardware and then discover there is no operator base, no procurement line item for the category, and no downstream analytics buyer. So you fly the missions yourself, process the data yourself, and deliver the insight yourself. If you build cars in a country with no drivers, you end up running a taxi fleet and calling it a product strategy. Every Indian deep tech hardware company became a services company, and the standard reading is that they failed to productise. It was closer to the opposite. It was an industry paying its own tuition, in cash, with no way to put the education on the balance sheet.

The Failure I Tell at Full Resolution

A book that makes this argument cannot afford to tell its successes better than its failures, so here is the failure that taught me the most.

We tried precision agriculture. Soil and crop testing at scale, roughly eleven lakh test results, distribution modelling across the sample, an advisory product designed on top of it. The global literature said this worked. Microsoft had case studies. Mahindra had applied research, including a study on tomato yield.

The assumption we started with, and the one that was wrong, was that Indian farmers were too poor to be a market. Southern landholding and income patterns complicate that badly. Drone ownership among farmers was already not trivial. The poverty framing was doing analytical work it could not support.

Then Reliance asked the question that reframed everything: could this reach every farmer in the country at roughly six hundred rupees a year?

The answer was no, and the reason was not the science, which worked. The optics infrastructure, the ground force, the last-mile agronomy layer, and the servicing economics did not exist at any price the arithmetic permitted.

Both consequences of the thesis fire at once here, which is what makes it the pivotal failure rather than merely a sad one. The blocks were missing on the supply side. And on the demand side, a farmer who received an accurate recommendation frequently could not act on it, because the input was unavailable locally, or unaffordable that week, or the credit cycle was wrong. We had built four layers ourselves in order to sell visibility to a buyer who could not act on what he saw.

The technology was ready and the market structure was not, and at the time we could not tell those two failures apart. Learning to tell them apart is most of what the next decade taught me.

Consequence Two: Visibility Is Not Actionability

The problem was never that the state did not know. Encroachment was known. Construction delay was known. Asset degradation was known. Crop stress was known. Tower faults were known, or knowable, long before anyone put a model on them.

Knowing changed nothing, because the money, the mandate, and the institutional path to act were absent. I spent years selling visibility to institutions that could see perfectly well and could not move. This one observation explains almost every failure I witnessed and every success. Where the capacity to act was already installed, deep tech worked immediately and unremarkably. Where it was not, the most sophisticated system in the country produced a dashboard nobody opened twice.

I learned this most clearly from the projects that broke the pattern.

We started, like everyone in the smart city and command-centre era, with drones as a 3D visualisation instrument. Surat proved that visualisation works. The models were good, the officials were impressed, the demonstrations landed. Then almost nothing happened, and working out why gave me the chain I have returned to ever since.

A 3D model is pointless. A 3D model with data attached is interesting. Neither matters until a decision is made on it. And a decision matters only if someone can fund the action it implies.

The fourth line is the one the industry took years to add. Most digital twin programmes stopped at the first. Stopping at the first is worse than never starting, because it consumes the budget and the political capital that the third and fourth lines would have needed, and it does so while producing something that looks, in a presentation, exactly like success.

Then COVID happened, and it was the one moment in the whole decade when the argument inverted.

The pandemic forced decision support systems into existence in months, systems no procurement cycle would have produced in years. They were statistical, forecast-driven, and simulation-based: combine every available dataset, estimate outcomes, let an official compare one strategy against another, and update continuously as the underlying reality moved. Federalism turned it into a natural experiment, because each state asked a different question. The northeast tracked commodities and foodgrain movement. Punjab built a system that changed shape as the crisis changed shape, from case tracking to mucormycosis as it emerged to live oxygen-tanker movement, and later to mobility estimation at telecom tower and pole level. Rajasthan went to individual geolocation, with all the governance questions that raises and that I am not going to pretend away.

Here is why it worked, and it is not what people assume. The technology did not improve. The teams were largely the same people. COVID was simply the one moment when the state's capacity to act briefly exceeded its capacity to see. Money moved without the usual friction. Authority was unambiguous. Inaction carried a political cost higher than action. In that condition, and only in that condition, visibility became instantly and obviously valuable, and the systems that delivered it were adopted in weeks rather than budget cycles.

What survived the pandemic was not the modelling. It was a competence I would put in the left column of the endowment table, built by accident under emergency conditions and genuinely rare internationally: crystallisation. The skill of compressing a large, messy, spatially distributed dataset into a single narrative that a specific person with specific authority can act on inside their decision window. That is the one thing the whole decade produced that cannot be imported.

When Does Seeing Become Doing

Once you have seen the inversion, you can generalise it. Visibility converts to action only when at least one of four conditions holds:

  1. A contractual or statutory consequence already attaches to the observed condition. Someone is obliged to respond, and the obligation predates the observation.
  2. The cost of acting is lower than the cost of the observed problem, and someone's budget line actually reflects that.
  3. Detection cost falls far enough that a previously uneconomic intervention becomes economic. Nothing new is learned, but the threshold moves.
  4. An external shock removes the usual friction on spending and authority. Rare, temporary, and not a business model.

Apply it backwards and the whole decade organises itself. Highway construction satisfied the first condition, which is why drone-based progress tracking worked there before it worked anywhere else. When we captured thirteen thousand kilometres of corridor, the reason it converted to action was not the quality of the capture. It was that the contract, the penalty clause, and the payment milestone already existed. The chain from observation to consequence was built, funded, and legally enforceable before we arrived.

Oil and gas satisfied the second. The NTPC tender, one of the first meaningful non-consumer AI procurements in the country, satisfied the third: the faults on those transmission assets were findable before the model, by inspection crews, and what the model changed was cost of detection per kilometre, and therefore the threshold at which acting became affordable. COVID satisfied the fourth.

Patent figure drawing of an inspection drone flying along power transmission towers.
FIG. 1 — Sensor payload pass along high-voltage transmission infrastructure.

Precision agriculture, most smart city digital twins, and most municipal deployments satisfied none of the four, which is why they produced elegant systems and no change whatsoever.

The prescriptive version, the thing I wish someone had told me in 2016: a govtech company's real diligence question is not whether it can detect the thing. It is whether the buyer can act on the detection, and who pays when they do. Almost no diligence process in the sector asks this, and it explains more variance in outcomes than technical capability does.

The AI That Was Already Everywhere

There is a myth that AI arrived in Indian government recently. It did not. Before anyone here used the phrase AI adoption, India already had enormous deployed computer vision. Number-plate recognition at tolls and city gantries. CCTV integration into command centres. Automated challan and speed enforcement. OCR across document-heavy workflows. Face recognition at national scale, with DigiYatra as the flagship.

All of it was narrow, extractive, and one-directional. AI was watching citizens and reading documents. It was not talking to anyone. And here is the pattern nobody noticed at the time: every single one of those deployments sat in a domain where the consequence of detection was already codified in law. A challan is actionability in its purest available form, because the observation and the consequence are the same object. India's early AI succeeded exactly where the first of my four conditions was already satisfied, and essentially nowhere else.

Around the same time, the industry did the healthiest thing it ever did, which at the time looked like retreat. Companies stopped trying to do everything and each chose one application. The drone-hardware companies went toward hardware and toward export, because domestic demand had a ceiling and the large contracts had already been distributed. The inspection companies went toward safety and asset integrity in heavy industry. The services firms went toward specific verticals rather than horizontal capability claims. Narrowing was correct. It just did not photograph like progress.

I should name the structural fact underneath all of this, the one that explains why India went drone-heavy in the first place. It was a substitution, not a preference. Countries with routine access to very high-resolution earth observation never had the LiDAR-versus-drone argument, never flew commercial-grade airframes past their design envelope, and never had to invent standard operating procedures in the field because their vendors had already written them. The absence of a satellite baseline handed India two things that belong in opposite columns of the endowment table: an unusually deep national competence in messy, mixed-source, ground-adjacent capture, and a decade spent solving an acquisition problem that comparable countries simply bought their way out of.

From Watching to Talking

The current wave is the first one where AI stops being a sensor and becomes an interface, and it is moving faster than anything before it. Understanding why requires both halves of the argument at once.

Take voice. Voice AI has moved from novelty to a recurring line item in government RFPs in a very short time. It is structurally correct for Indian public service delivery: multilingual by necessity rather than as a feature, low dependence on literacy, no app to install, working on the phone people already own over the network they already have.

Now notice why it moved so fast, because it is the strongest evidence this whole framework has. Two independent arguments, developed separately across everything above, point at the same category. Linguistic diversity is an unusual Indian input, so the endowment argument predicts India should lead here rather than follow. And voice is the first AI category in Indian govtech where the system's output is itself the action, so the actionability rule predicts the conversion problem that killed a decade of projects simply does not arise. When two arguments built on different foundations converge on the same prediction, and the prediction holds, that is about as much as an argument of this kind can hope to demonstrate.

Video is the second front. For a country this large, this dense, and this varied, continuous visual sensing is the only mechanism that can plausibly close the gap between what is happening on the ground and what any central authority knows. Nothing else scales to the geography. What separates it from surveillance theatre is a set of conditions, not a set of reassurances: geotagged integration so an observation is anchored to a place rather than a person, ground-verification loops so model output is checked rather than trusted, and an audit trail from observation to decision so that when action follows, its basis is reconstructible. The governance question and the utility question turn out to have the same answer. A video system that cannot show its chain from observation to decision is both illegitimate and useless, because an official who cannot defend the basis of an action will not take it.

The third front is the channel, and it is where I want to sound a warning rather than a celebration. Citizen engagement is collapsing into WhatsApp, the channel people already live in. The obvious benefit is reach. The less obvious one is that agentic chat layers change the government's own epistemic position: engagement and grievance data arrive structured, continuous, and attributable, rather than as a quarterly report assembled by the same people being evaluated by it. For the first time the state can see its own service delivery close to real time.

A grievance channel that outpaces the redressal capacity behind it does not close the gap between seeing and doing. It industrialises the gap, scales it, and makes it visible to every citizen who uses it. The failure mode of this generation of govtech AI is not that it will not work. It is that it will work perfectly at the front end and expose exactly how little capacity sits behind it.

The previous generation of dashboards was visible only to officials. This one is not, which makes the failure mode politically live in a way it never was before. This is the point at which my central argument stops being retrospective and becomes something closer to a caution.

The Blocks Arrive

For the whole of the period I have described, the defining fact was that the blocks did not exist. That is now changing, and the last part of the story is whether the change is real.

The semiconductor push is the headline. There is a mission, an incentive structure, fab commitments, packaging and assembly capacity, and a design ecosystem that was always stronger than the manufacturing one. But a fab does not help a drone company next year, and pretending otherwise is how the sector loses the credibility of the people it most needs. The honest sequence from announcement to a qualified domestic part that a product company can actually design into a bill of materials is measured in years and passes through several steps that receive no coverage. The open question, and I want to leave it open, is whether it arrives as a genuine supplier tier or as a set of facilities without an ecosystem around them. That second outcome is the Make in India failure mode repeating at roughly a hundred times the capital intensity.

But silicon is not actually the binding constraint for the sector this essay is about, and this is the prescription the whole argument has been building toward. Every company I described in the hard years was blocked by components, not by silicon, and none of them was blocked by the absence of a frontier research lab. Sensors, optics, actuators, power electronics, precision machining, connectors, test and qualification infrastructure: this is the tier that determines whether Indian deep tech hardware becomes viable, and it receives a fraction of the attention because it produces no announcement.

This is where the endowment argument earns its keep by turning into a recommendation. Compute and research depth are absent inputs that take a generation to create, and complaining about them is close to complaining about the weather. Components are absent inputs that a decade of deliberate policy and patient capital could have created, and could still. That is a harder criticism than the usual one, precisely because someone could act on it.

There is also a category of problem that no other country will solve for India, because it arises from exactly the endowments that make India distinctive. Algorithmic decisioning in an economy where most transactions are undocumented, so absence of record is the norm rather than the signal it is treated as elsewhere. Federal capacity asymmetry, where the same system lands in states with wildly different ability to operate, audit, and correct it, so a safeguard that is real in one state is decorative in another. Language coverage as a question of who receives the service and who does not, once the interface is voice. The boundary between national sensing and surveillance, in a country with this density of cameras and this thinness of settled precedent. These are not reasons to move slowly. They are the specification. A country whose unusual endowment is scale and heterogeneity has a governance problem shaped exactly like its opportunity, and any framework imported wholesale from a jurisdiction with different endowments will be either unenforceable or irrelevant.

Is India Behind?

Now the claim everyone actually wants adjudicated, which I have deliberately withheld until the framework could carry it.

The claim in circulation is that this is the first technology wave in which Indian government adoption trails the private sector by a small margin rather than a generational one. Test it layer by layer, and be genuinely willing to falsify it.

It holds for voice, vision, document AI, and conversational delivery. Every one of these is a category where India's inputs are unusual and where the action is either already codified or is itself the output. Both halves of the argument predict convergence there, and the evidence shows it. It does not hold for model development, compute, talent retention, data infrastructure, or procurement speed, where the lag is real and in some dimensions widening.

What explains the difference is not effort, not ambition, and not the quality of officials or founders. It is endowment, and the presence or absence of a path from observation to action. Which means the convergence is real and considerably narrower than the optimistic reading suggests. India's government is not catching up with the frontier. It is competitive at the specific layers where the country's inputs are unusual. That is a smaller and more defensible claim than the one currently being made, and a more useful one, because it tells you where to build.

I will end on the uncomfortable question, since I promised to. The founder generation I belong to was selected for endurance, and that was the correct training for a decade whose binding constraint was survival with no adjacency to lean on. The current cycle rewards the opposite disposition. It is an open question, and not a rhetorical one, whether the people who won the last round by outlasting everyone are suited to a round decided by speed. I do not know the answer. I am one of the people the question is about.

Back to the Dam

Return to Polavaram. Return to the woman and the line about the villages.

Ask what would be different if the same project started today, and the honest answer has two halves. The first half is genuinely better. Better base data, faster capture, a voice channel that could reach every affected household in its own language, a grievance system that would register every claim within a day. Everything on the seeing side has improved, some of it dramatically, and a good deal of that improvement passed through my own hands.

The second half is longer, and the sector avoids it. The compensation would still be contested. The record would still disagree with the ground. The money to act would still arrive on a schedule set by something other than the evidence. Almost nothing on the doing side has improved.

That asymmetry is the whole decade in one image, and it is the reason the next one is a policy problem wearing the costume of a technology problem. We got very good at seeing a country that we are still, institutionally, only barely able to act upon. The blocks are finally arriving. Whether we can act is a different question, and it was always the harder one.

Footnotes

  1. Polavaram multi-purpose project on Godavari river, considered the lifeline of the State, was first proposed in 1941 and preliminary investigations were carried out between 1942 and 44. The New Indian Express, Proposed in 1941, but still a dream (2017-11-05). Establishing that the Polavaram project was first formally proposed in 1941 substantiates the author's framing that the project had been in the planning stages for roughly fifty years prior to active on-site work around the turn of the 21st century.