I would know that Digital-World Urbanization (DWU) is wrong if its constructs explained nothing that an enterprise architecture maturity score doesn’t already explain, if they turned out to be indistinguishable from instruments that already exist, or if two analysts couldn’t classify the pieces of the same company the same way. Those are the three discriminant tests I wrote into the preprint, and this post works through them because a framework that can’t fail can’t explain anything either.
I’m the worst-placed person to judge this theory, because I wrote it. DWU proposes reading a company’s digital world as an inhabited territory organized into seven systems (the enterprise as an inhabited territory), and I’ve spent eleven posts of the Digital urbanization series using it. This one puts in writing what would have to happen for me to stop using it, or to use a lot less of it.
The actual status of the work
DWU currently exists as a conceptual preprint on Zenodo, with no peer review. Its contributions are theory-building claims, and none of them is an empirical finding. Its six propositions are hypotheses nobody has tested yet.
It rests on a preliminary integrative review of 28 references published between 1977 and July 2026. The review was designed to build constructs and makes no claim to be exhaustive. Its methodological limitations are declared in the text itself and I won’t soften them. A single author designed the search, coded the sources and chose the constructs. The protocol wasn’t pre-registered. Database coverage is selective, the corpus is mixed, and the bibliographic audit didn’t independently assess the quality of each source. There may be interpretive bias, and traditions written in languages other than English and French may be underrepresented.
Then there’s the part that bothers me most. During the three years before I wrote the framework down, I applied and refined these ideas in enterprise platform architecture work. Of that practice, only the part done on Grupo Diveco’s corporate platform, which starts in August 2025, is documented in repositories. All of it plays an abductive role, meaning it suggested the constructs and shaped their boundaries. Nobody designed it as a study and it validates nothing. It does carry a risk of confirmation bias, because the practice the ideas came from is also where I see them confirmed.
When a point-to-point integration breaks, I read it as an improvised road. When a team builds its own notification service, I see public infrastructure that didn’t arrive in time. Those readings can be right and still be weak evidence, because I make them with the category already in place.
Portal Diveco, also called Ciudad Diveco internally, is Grupo Diveco’s internal corporate portal and the documented part of that practice. As of September 2026 it brings together some fifty registered tools, which run on shared identity, audit and row-level filtering services and on a data lake with shared semantics. It also has an AI agent whose scope is enforced server-side from the user’s identity. In September we also published a 3D model that draws the platform as a city, one building per tool, with static data synced by hand. The framework was abstracted from that practice, and the portal doesn’t use its vocabulary. The urban language is in the brand, like the “Ciudadano DIVECO” role every user has received by default since August 2025.
All of this shows that the constructs can be operationalized in a real organization. It doesn’t validate any of the six propositions. It’s a single case, I’m its architect, and almost all of it predates my writing the framework, since it’s precisely the practice the constructs came from. I also have no comparative measurements of duplication, integration fragility, authorization incidents or adoption. What it offers is material for the retrospective case study where validation would start, if the organization authorizes it.
The series has the same bias on a small scale. Northstar Consumer Group, the fictional case that runs through it, is something I composed from patterns in my practice that the literature documents, so that every system would have a visible problem, and I chose every example because it illustrated a system well. None of them shows a problem where the framework has nothing useful to say. A case written so the framework can explain it can’t prove that the framework explains anything, although it’s useful for rehearsing a protocol.
Rival explanations
Suppose that a year after adopting the DWU vocabulary, Northstar replaces its three notification implementations with a single shared service. I’d read that as the reduction in duplication that proposition P1 attributes to civic infrastructure. But a new CTO might have arrived and pushed the consolidation through, or the platform team’s budget might finally have been approved. The company might have grown until three services became unsustainable, or a regulator might have demanded an audit trail for every message sent to customers. The architecture practice might have matured, or two very good platform engineers might have joined.
Executive sponsorship, platform funding, scale, regulatory pressure, architecture maturity and talent are the six rival explanations the framework acknowledges. None of them needs districts or a subsurface, and if any of them explains the change just as well, DWU has put a new name on something already understood. Sponsorship and talent worry me more than the rest, because they can sit behind almost any improvement without leaving a trace in an inventory.
Habitability has a rival of its own, which the framework also records. A platform can win adoption because it’s easy to inhabit, as I argue in digital habitability, or because its use is mandatory, because incentives are tied to it, or because someone with power decided it. On a usage dashboard, those adoptions look the same.
In building evidence before consensus I argued for prototyping the reality you want to argue for, and here that stance shows its limit. A working prototype proves something is possible. Claiming that DWU explains something also requires ruling out the other six stories.
Three tests that could bring it down
The three discriminant tests share a shape. Each one sets a DWU construct against something that already exists and fixes in advance which result would leave it with no reason to exist.
The framework is judged by four criteria, which are explanatory utility, conceptual distinctiveness, operationalizability and empirical validity. I read the three tests as the operational form of the first three. Empirical validity is left to testing the six propositions.
Incremental explanatory power
District jurisdiction, mobility governance, informal settlement density and habitability would have to explain differences between organizations in duplication, integration fragility or shadow-tool adoption, beyond what enterprise architecture maturity and the other controls the framework lists (platform capability, IT governance centralization, digital maturity, organizational size) already explain. In companies like Northstar, where the BI tool and one country application read ERP tables directly, the question would be whether integration fragility (what breaks when one of those tables changes) is better predicted by a measure of mobility governance than by a maturity score. If both predict the same thing, DWU is a maturity assessment with urban labels.
Construct distinctiveness
Digital habitability should correlate only moderately with generic user satisfaction or technology acceptance instruments, and informal settlement density should separate from conventional technical debt inventories. Northstar’s adjusted-price spreadsheet, the one that circulates by email, makes me expect that separation, because it doesn’t live in any repository and a technical debt inventory built by scanning code would be unlikely to find it. An example I picked myself proves nothing, of course. If measures in real organizations converge above accepted thresholds, the constructs are redundant.
Inter-rater reliability
Independent analysts classifying the same company into the seven systems should reach substantial agreement. If disagreement persists, the urban vocabulary lacks the stability that cumulative research needs, because each study would be measuring different things under the same names.
Of the three, it’s the only one I can start running without data from any organization.
An experiment with Northstar
Two analysts who had no part in writing the framework receive Northstar’s component inventory and the definitions of the seven systems exactly as published, and each one classifies on their own. An illustrative draft of the protocol looks like this.
# Illustrative: pilot inter-rater classification round on the fictional
# Northstar case. Not a validated instrument or a standard.
round: northstar-pilot-1
raters: 2 # had no part in writing the framework
communication_during_round: none
definitions: rodas-lopez-2026 # the paper's, with no glosses from the author
criterion: current_state # what the component is today, not what it
# would be after an intervention
assignment:
primary_system: required # agreement is computed on this field
secondary_system: optional
justification: one_line
components:
- erp
- ecommerce-notifications
- svc-integracion
- integration-policy-check
- purchasing-agent
- supplier-master-data
- purchasing-policy-entity-a
- adjusted-price-spreadsheet
- vector-index # only if some agent retrieves documents
preregistration:
agreement_statistic: fixed_before_the_round
substantial_agreement_threshold: fixed_before_the_roundI’d expect quick agreement on pieces like the ERP (a building), Finance (a district) or the five legal entities (territory). The predictable disagreement sits on the borders the diagram marks.
DWU lists vector indexes under the subsurface, but if one of the twelve experimental agents queries an index built by its own team, that index has an owner, a purpose, users and a lifecycle, which is how the framework describes a building. Policy engines appear as public infrastructure, yet one of the policy checks Northstar keeps rebuilding, if it lives inside an integration, also reads as a traffic rule, which puts it in mobility. The svc-integracion account is a non-human inhabitant by definition, even though it exists only to move data. One entity’s purchasing rules are jurisdiction or a Supply chain decision depending on who reads them, and supplier master data leaves the same doubt between subsurface and building.
The purchasing agent is the case I find most interesting. As an actor that prepares purchase orders, it’s an inhabitant, even though today it runs on credentials inherited from a developer. As a deployed component with a model and an interface, it’s a building, and DWU’s diagnostic table places agent interfaces among the buildings.
The first risk is that some of the disagreement will come from the protocol. The e-commerce notifications are a building if you classify current state and public infrastructure if you classify function, which is why the draft fixes criterion: current_state. Granularity matters too. The definitions put operational records in the subsurface, and the agreement I predict for the ERP assumes nobody separates the application from its tables. If it took many rules like that to get two analysts to agree, that would already be a result.
The second risk is scope. A round on Northstar only rehearses the protocol. The real test needs inventories from real organizations and analysts who haven’t read this series, and there it will matter as much where the disagreement concentrates as how much of it there is. If it piles up on one or two borders, those borders need redefining. If it spreads across almost every pair of systems, the whole vocabulary is the problem.
What would make me abandon parts of the framework
If the framework fails on any of the four criteria, the answer is to narrow it, revise it or abandon parts of it, and keeping the metaphor for rhetorical reasons is ruled out in advance. I’m writing down which result would lead to which cut, so I can be held to it.
- If, on real inventories, disagreement between analysts concentrates on the border between subsurface and buildings, I’d merge those two systems or narrow the subsurface to shared data with its own owner.
- If digital habitability converges with satisfaction or technology acceptance instruments, I’d retire the construct and use those instruments.
- If informal settlement density doesn’t separate from inventoried technical debt, the construct leaves the theory, even if the four responses (recognition, integration, relocation and retirement) remain useful in practice.
- If district jurisdiction and mobility governance explain nothing beyond architecture maturity, I’d present DWU as a communication vocabulary for architects and stop calling it a theory.
- If a retrospective study of my documented practice doesn’t find in the decision records and inventories the traces my memory attributes to these ideas, I’d treat that practice as anecdote.
Even if it passed all three tests, the framework has limits I already know about. The most serious is overextending the analogy, which can cover up phenomena better explained through ecosystems, networks, markets, institutions or biological metaphors; an analogy you’re fond of tends to absorb them all. The seven systems are also provisional. Constructs may be missing (the framework itself names the digital economy, environmental sustainability, power relations, institutional legitimacy and temporal rhythms), and others may turn out to be redundant or merge.
At Northstar, the most visible gap is power. The recurring temptation to buy one vendor product per district and call the license org chart an architecture is a budget phenomenon, because whoever controls purchasing ends up drawing the districts. DWU names the symptom and has no construct that explains why it happens. Legitimacy is the same story. A shared identity service can be habitable and still go unused because the team that runs it has no recognized authority over the districts.
That leaves scope. DWU assumes a medium or large company with the authority to set common rules, and its assumptions about agents will need checking against emerging standards (Booth et al., 2026), because the concept is moving fast.
Where validation would start
DWU proposes validating in four stages, each more expensive than the one before.
- Construct validation with expert panels.
- Design-science development and evaluation of diagnostic artifacts (building the artifact and evaluating it in use), for example comparing the decisions produced by the ten-question canvas with those produced by an application inventory, a capability map or an AI risk checklist.
- Comparative case studies across organizations.
- Quantitative tests of the six propositions.
Two preliminary steps come before those stages. A reproducible systematic review addresses the weakness of the current review protocol. The other step doesn’t depend on me alone. It’s a retrospective case study of the three years of practice, based on documented evidence (architecture decision records, integration inventories, cycle time measurements, incident logs). Grupo Diveco is the candidate, since it already holds part of that evidence (git history, an ADR, technical findings documents and an audit log), though only from August 2025 onward. It’s conditional on organizational authorization, and that condition is real. If authorization doesn’t come, the route starts at the expert panels.
Even with authorization, that study carries the problem of the practice it reviews, since I would be doing it on decisions I made. I’d add a second reviewer to code the same documents without knowing my conclusions.
That second reviewer is one instance of the rule I apply against author bias in every study I’m proposing, which is to separate roles. People who didn’t write the framework do the classifying and coding, thresholds are fixed before anyone sees data, and rival explanations enter as controls from the design stage. That has a cost. Validation is slow, it depends on permissions I don’t control, and meanwhile I keep using the framework in my work, which feeds the very bias I’m trying to contain. I don’t have a clean way out of that, beyond stating it.
The test that worries me most is the third, because it’s the cheapest. An architecture team can run it with its own inventory, two people who haven’t read this series and the published definitions of the seven systems, without asking anyone’s permission. If someone does, and the disagreement lands where I didn’t predict it, that result will teach me more than the eleven posts before this one.
This article is part of the Digital urbanization series, based on the preprint Toward a Theory of Digital-World Urbanization (Rodas López, 2026), a conceptual framework whose ideas I applied and refined on Grupo Diveco’s corporate platform, though its six propositions have not yet been empirically validated. Previous: The digital urbanization canvas: 10 diagnostic questions.