← All papers
Thesis Technology · Society · Governance 2026

The Limits of Data Integration in AI Decision-Making Systems

Why the ethical failures of AI data platforms are built into their architecture, not patched on after.

Abstract

This thesis argues that the ethical problems associated with large-scale AI data integration systems are not technical defects to be patched after deployment, but emergent properties of the underlying data architecture. Existing scholarship examines either the institutional setting (surveillance studies) or the algorithm itself (AI ethics), leaving the architectural middle unexamined: the layer where decisions about integration, classification, storage, processing, and visualisation are made. Drawing on Science and Technology Studies, together with three layer-specific extensions (Actor-Network Theory, Verbeek's technological mediation, and Nissenbaum's contextual integrity), the thesis asks how design decisions made before, beneath, and around the algorithm produce the patterns observers later name as ethical issues. The approach is qualitative and design-oriented, combining a single-case study of Palantir Foundry, a conceptual mapping of six architectural layers (idea, collection, integration, storage, processing, visualisation), and a purpose-built design probe, the Noor System, a footfall-tracking platform that reproduces Foundry's structure at a scale where every decision is traceable. The analysis is organised around four mechanisms: more data does not equal better outcomes; AI as amplifier, not origin; classification and definition power; and friction as architectural intervention. Together they show that consequences belong to no single layer but arise from the interaction of all of them, which is why a better algorithm cannot remove them and a component-level audit cannot locate them. The thesis concludes that intervention must therefore be architectural, with deliberate friction at the layers where consequences are produced, and that such friction is not a substitute for regulation but its precondition.

Key words: data integration platforms; emergence; Science and Technology Studies; Palantir Foundry; design probe; architectural friction

1. Introduction

1.1 The rise of AI-driven data integration

The predominant institutional response to problems of coordination and control has been large-scale, AI-driven data integration platforms. Lyon (2022) maps the reorganisation of surveillance from observation to dataveillance as infrastructural rather than episodic. This thesis aims to examine this shift. Contemporary surveillance is not just more observation, but an emergent property of integrated data architecture. It exists even when nobody is watching, and it infers behaviour not present in any single record. Neither capacity is a property of surveillance; each occurs only after data is integrated.

These are not theoretical platforms. Documented deployments of Palantir, by the Los Angeles Police Department (Brayne, 2017), by the U.S. Immigration and Customs Enforcement (Knight & Gekker, 2020), and across European public-security contexts (Ulbricht & Egbert, 2024), demonstrate a tightly coupled architecture of integration, classification, and analysis (Iliadis & Acker, 2022), operating where its outputs determine whether individuals are stopped, deported, or otherwise rendered visible to state power.

1.2 The problem with current framings

There are two strands of response to this discussion and both miss the same point. The first approach sees ethical issues as products of algorithms that can be addressed post hoc by normative concepts like fairness, accountability, transparency, etc. (see Mittelstadt et al., 2016, for the standard taxonomy). The second sees them as engineering issues, addressed with improved training sets and post-hoc bias audits. Pasquale (2015) demonstrates that this is not enough: opacity is not an incidental deficiency of an algorithm, but rather a structural characteristic of integrated systems, an "artifact" of the underlying architecture and not of the model.

The two sides have one assumption in common: ethics can be applied to a system post-release. They treat the platform as a completed structure, first built, then inspected afterwards, with safeguards added as rooms are added to a finished house. This thesis argues the opposite view: ethical character cannot be added onto a data architecture but grows from within it. Its conditions are chosen, which is what keeps its consequences a matter of design rather than evolutionary fate. The issue is not with any one component but with the interaction between the components once they are assembled. This layer of data architecture itself is the unit of analysis which has not been the subject of either surveillance studies or AI ethics.

1.3 Research question and central claim

This thesis asks: how do data integration and design decisions in large-scale AI systems influence outputs and produce downstream consequences that observers later interpret as ethical issues? The phrasing is causal-architectural, not evaluative. It does not ask whether such systems are ethical, or how to make them ethical. It asks how decisions made before, beneath, and around the algorithm produce the patterns observers later call ethical issues. The claim follows. What we describe as ethical issues in AI data integration systems are a byproduct of the architecture. Fixing them requires architectural change, not technical fixes alone. The argument is not against integration platforms. It is about how they should be built. Rejection is the easier position. Redesign is the harder one.

The term emergent is used with intent and precision. For complex systems, a property is emergent if it is not possessed by, nor predicted by, any particular component of the system in isolation, but is instead a product of their interactions in accordance with a specific set of rules (Holland, 1998). The ethical consequences of a data architecture are fully grounded in that architecture. Nothing mystical is claimed. Yet they cannot be read off its design specifications in advance; they become visible only once the architecture is assembled and run; this is the weak form of emergence used in the thesis. The concept is not only framing but used to explain what no other term could. §6.3 demonstrates that an ethically designed system, deployed fifty times, aggregates into a whole with a property no single deployment has.

Emergence in this sense is the spine of the thesis. The theoretical tools, the six architectural layers, as well as the design probe are all instruments to locate where, and how, the ethical consequences emerge. The thesis is situated in the field of Science and Technology Studies; it builds on Actor-Network Theory, technological mediation by Verbeek, and contextual integrity by Nissenbaum as extensions of the STS condition that technology is not neutral. In developing the claim, it adopts Palantir Foundry as its empirical reference, and a design probe based on a purpose-built footfall-tracking system, the Noor System, which replicates Foundry's six-layer structure at a scale where all architectural decisions are traceable. Palantir's own self-presentation is critically read across its governance and commercial registers, the contrast between them providing further evidence of the way in which architectural choices can be framed for different audiences.

1.4 Academic and societal relevance

Academically, the thesis offers a contribution to the field of critical data studies, defined as the study of data assemblages as opposed to data as neutral information (Iliadis & Russo, 2016) and connects it with the older STS tradition, based on Winner's (1980) claim that artifacts are distributions of power, and the infrastructure-studies perspective on classification systems as moral and political infrastructures (Bowker & Star, 1999). The contribution is the bridging: the positioning of ethics upstream, within architecture where consequences are created, not downstream in outputs or interpretations. The thesis is neither for nor against integration systems. The focus is on how these platforms should be built and why.

On a societal level, the stakes are not hypothetical. Palantir deployments have been documented to lead to misidentification in enforcement databases, over-policing of already datafied populations, and accountability gaps across institutions and national borders (Brayne, 2017; Ulbricht & Egbert, 2024). Foundry is the chosen case because it is among the most documented and the architectural reading created here is meant to be transferred. Browne (2015) provides the historical depth, exploring how the long-established practice of racialised social sorting facilitated contemporary surveillance. Furthermore, the populations rendered most visible to integrated surveillance systems are the ones who were always most observable. Naming the architectural locus of these consequences is the first step for any intervention.

1.5 Structure of the thesis

The thesis is organised into seven chapters. The architectural gap is identified in Chapter 2 through a review of the literature on surveillance, AI ethics, Palantir and critical data studies. The theoretical framework is developed in Chapter 3, with STS as the major lens and Actor-Network Theory, mediation theory, and contextual integrity as layers in specific places. The methods used in the qualitative, design-oriented study are presented in Chapter 4: the case-study analysis of Foundry, the conceptual mapping of the six layers, and the design probe. Chapter 5 discusses the analysis around four types of mechanisms: more data does not equal better outcomes; AI as amplifier, not origin; classification and definition power; and friction as architectural intervention, based on the empirical record and Palantir's own presentation. Chapter 6 outlines methodological, scope, ethical, and reflexive limitations. Chapter 7 revisits the main thesis and maintains that these systems are not to be discarded but to be restructured.

2. Literature Review

The chapter contextualises the thesis in four fields of scholarship: surveillance studies, AI ethics, empirical studies on Palantir, and critical data studies. Each diagnoses contemporary data infrastructures at a different layer. Read together, they provide a detailed description of the algorithm and the institutional context but fail to analyse the architectural layer between them as a unit of analysis.

2.1 Surveillance, policing, and the reshaping of institutional practice

Surveillance studies shape the environment in which integration platforms are used. Lyon (2022) outlines the concept in four phases: observation, sorting, digitisation, and dataveillance, to argue that surveillance is now infrastructural rather than discrete. Gandy (1993) was among early critics who predicted the move: he viewed the panoptic sort as a systemic characteristic of integrated personal-information systems, which merge administrative records and sort people regardless of whether a rule is violated. Brayne (2017) continues exploring this idea in modern police work, as the LAPD's use of Palantir transformed subjective risk assessment into risk scores and reduced the standards for inclusion in the database. As previously mentioned, Browne (2015) provides the historical and racial background about surveillance, showing its shift in use from slave passes to biometrics. Together these works reveal that the consequences of data integration platforms are not distinct from but rather an extension of surveillance practices that preceded them.

2.2 Two dominant strands in AI ethics: normative and technical

AI ethics is split into two directions, both of which consider ethical issues as downstream of architecture. The first is a normative one: Mittelstadt et al. (2016) introduce the canonical taxonomy, which classifies algorithmic ethical issues into six categories: inconclusive evidence, inscrutable evidence, misguided evidence, unfair outcomes, transformative effects, and traceability, each to be audited once a system operates. On this view, ethics is applied to the system, not embedded in it. The second strand is technical, in which ethical issues are approached as bugs to be patched with improved training data, fairness-enforced models, and post-hoc audits. That is not structurally enough, as Pasquale (2015) demonstrates: opacity is the product of the architecture supporting the model, not the model itself, and the problem is upstream of the model. Zuboff (2019) extends the critique to political economy. She names surveillance capitalism, an economic system that treats people's lived experience as unclaimed raw material to be harvested and turned into forecasts of what they will do, and a novel power she calls instrumentarianism, which both reads and steers behaviour toward someone else's goals, working through pervasive networked computing rather than force. Behavioural surplus is fabricated into prediction products and traded in what she terms behavioural futures markets. Her account names the power and its economic logic but stops at the market form; it does not reach the architectural level at which those consequences begin, which is exactly what this reading supplies.

2.3 Empirical research on Palantir and integrated data platforms

Empirical literature on Palantir is limited but strong. The key peer-reviewed work is by Iliadis and Acker (2022), who reverse-engineer Palantir's stack from 155 patents and primary documents, and find that the dynamic ontology is a customisable per-client classification system, that heterogeneous sources are integrated in a three-stage data funnel, and that forward-deployed engineers co-construct customer architecture. Knight and Gekker (2020) analyse Palantir's Investigative Case Management (ICM) platform. They reveal that it is the logic of knitting together twenty-one databases for U.S. Immigration and Customs Enforcement that produces the harms usually associated with algorithmic decision-making, not the AI sitting on top of it. Furthermore, Munn (2017) provides the reading on the visualisation layer and elaborates how Palantir Gotham's interface anticipates and preconfigures suspicious patterns, mediating the relationships operators come to see. Ulbricht and Egbert (2024) bring the record to Europe and pinpoint the challenges of effective platform regulation in the public-security space.

This record is complemented by Palantir's own self-presentation in the thesis. In addition to the privacy and governance whitepaper (Palantir Technologies, 2024), Chapter 5 discusses two AIP product pages: one for Export Control Compliance and another for Procurement Fraud Detection (Palantir Technologies, n.d.-a, n.d.-b) to illustrate how Palantir markets its architecture. Where there is a tension between efficiency and privacy in the whitepaper, this contrast disappears in the commercial pages and is instead used as a rhetorical argument in Chapter 5.

2.4 Critical data studies and the construction of data

Critical data studies (Iliadis & Russo, 2016), which focus on data assemblages as opposed to data as neutral information, provide the vocabulary for rendering data architectures as subject matter for critique. In his edited volume, Gitelman (2013) claims that "raw data" is an oxymoron, as data are always pre-classified, pre-selected, pre-formatted. Kitchin (2014) expands this to a critique of the "end of theory" empiricism on which integration platforms implicitly rely, distinguishing data-driven from knowledge-driven science, and boyd and Crawford (2012) provide six provocations on big data, chief among them the notion that more data does not equal better knowledge and that claims to objectivity in big-data analysis are misleading. These works deny the possibility of using data as a straightforward input to computation: data are already always structured, and structure is a political and epistemological matter.

2.5 The architectural gap: what existing literature does not address

Each tradition advances on a different level: surveillance studies on the institutional setting, AI ethics on the algorithm, Palantir-specific studies on specific deployments, and critical data studies on the constructed nature of data. None explores the architectural middle, the level at which decisions about integration, classification, storage, processing, and visualisation determine what a system can later be made to do. Singh et al. (2018) contend that accountability issues stem from data-flow architecture in interconnected systems-of-systems, and Cobbe et al. (2023) extend this to the "accountability horizon", the structural end of visibility that creates accountability gaps in algorithmic supply chains. Stahl (2025) introduces the reification problem, noting that categories become things instead of tools. At the governance level, Koskinen et al. (2023) suggest "house rules" as a way to create "procedural friction". All identify the problem as architectural but approach it through one discipline and at one level.

There is one precedent that goes back further. Perrow's (1984/1999) Normal Accident Theory posited that catastrophic failures of tightly coupled and interactively complex systems like nuclear plants, chemical plants, and aviation were accidents of the system, meaning that they were not necessarily caused by any single element of the system, but by the nature of the interaction between elements, and that they could not be "fixed" simply by making better parts. Perrow's domain is physical safety, and the transposition is intentional: what Normal Accident Theory did for industrial catastrophe, an architectural reading of integration platforms can do for the consequences observers call ethical, locating them in the interaction of layers rather than in any layer alone.

The contribution of this thesis is in reading the architecture as a whole, across the six layers and four mechanisms, as a sociotechnical object where the ethical consequences are manifest through the interactions between the parts.

3. Theoretical Framework

This chapter describes the analytical perspective used in the architectural reading. Science and Technology Studies offers that lens and three other traditions: Actor-Network Theory, Verbeek's technological mediation, and Nissenbaum's contextual integrity, extend it at specific architectural layers, while a cross-cutting framework on classification, inspired by Bowker and Star, runs through them all.

A reader coming across four theoretical reference points in a thesis may rightfully ask whether the analytical angle has been settled. The objection deserves to be answered in itself, and because it provides a clarification of method. These are not four competing frameworks within the thesis but one framework (Science and Technology Studies) and three specifications. STS offers the premise that technology is not neutral, and a premise is not an analytical tool: "technology is not neutral" indicates the place to search but not what to search for. The thesis therefore requires tools fine enough to dissect specific layers, and more than one is required because of the object of study. Large-scale integration platforms are not one type of system: they combine police information, immigration data, financial data, and social-media data into a single analysis environment, and the impact of this combination occurs in various layers and forms (Knight & Gekker, 2020; Iliadis & Acker, 2022). Actor-Network Theory is about the distribution of agency in the integration network; Verbeek's mediation is about what the visualisation layer produces as an image of what operators see; Nissenbaum's contextual integrity is about what the integration layer breaks. None of them goes to all three sites. The diversity of the architectures is reflected in the diversity of the theories, not vice versa. Each is used where it is effective and the others fail. The four tools are used to track the process, and the process being tracked is emergence.

3.1 Technology is not neutral: the shared premise

The underlying assumption is that technology is not neutral. Winner's (1980) canonical argument about the Long Island bridges exemplifies how technical arrangements can have a political nature, even when they are not deliberately designed to be such. Artifacts carry and perpetuate specific configurations of power and design decisions can foreclose democratic deliberation in a manner that is difficult to correct. It is the foundational claim of Science and Technology Studies. This is where the organising concept of the thesis starts. When technology is not neutral, the non-neutrality of one feature, a data field, a category, an interface option, is hardly significant in isolation. When coupled with the non-neutrality of all other components, it acquires meaning and becomes a cause of a system-level effect that no individual component would have on its own. This is the sense in which Holland (1998) speaks of emergence, and it is why the thesis can only read the architecture whole, but not audit its parts: an emergent property cannot be found by auditing each part individually.

3.2 Science and Technology Studies as the analytical container

Science and Technology Studies is the field within which the other three theories sit. Its common commitments, as summarised by Sismondo (2010), are as follows: the critique of technological determinism, the understanding that technical arrangements are influenced by social, political, and economic forces at the time of their design, and the principle of co-production between technical and social orders. Bijker et al. (1987) established the Social Construction of Technology (SCOT) programme, showing how social groups shape technological development through negotiation and "closure". The thesis uses this to argue that architectural choices, too, are negotiated. Two additional sources help frame integrated data platforms. Star (1999) lists nine properties of infrastructure; four are most relevant here: embeddedness, reach, embodiment of standards, and becoming visible upon breakdown. These properties justify treating Palantir Foundry as infrastructure rather than software. Plantin et al. (2018) state that modern digital systems are both platform-like and infrastructural. Foundry is both. As a commercial platform it carries vendor lock-in and value-extraction logics. As an infrastructure its presence reorganises institutional practice.

3.3 Actor-Network Theory and distributed agency

Actor-Network Theory, within the STS frame, offers a specification for non-neutrality of technology in practice. The basic assertion of Latour (1992) is that mundane artifacts carry morality and politics in the script that they impose on users. Latour's seat-belt buzzer, which enforces a behaviour that the human agent alone might neglect, demonstrates that artifacts can be analysed as actors rather than inert tools. Callon (1986), drawing on the scallop fishermen of St Brieuc Bay, supplies the vocabulary of translation, enrolment, and obligatory passage points. These terms describe how networks come to speak through a single point, how actors come to be captured into roles. Law (1992) extends this with the concept of heterogeneous network ordering, which is the refusal to accept the existence of a separate social sphere from technical and material relations. There are two things that ANT does for this thesis. First, it means treating Palantir's data fields, ontology objects, and integration logic as actors that exert agency alongside the analysts and institutions deploying them. Second, it describes how Foundry orchestrates its data-flow networks. Foundry translates vast amounts of heterogeneous data into a common format, enrols actors into a single network, and becomes the obligatory passage point through which institutional decisions must pass.

3.4 Technological mediation and the visualisation layer

Informed by the STS framework, Verbeek (2011) elaborates the theory of technological mediation, moving ethics from something applied to something practised through technology. Artifacts do not neutrally represent reality. They construct what an operator can see, and through that, what they can do. The framework is based on post-phenomenology, which treats the moral dimension of design as a property of the artifact itself. When Palantir Gotham draws network diagrams or colours some entities red and others green, these are mediation effects at the visualisation layer. The chart type and its framing construct what operators see and act upon. While Munn (2017) provides the empirical reading of how the interface prefigures suspicious patterns in Gotham, Verbeek gives that layer its name; the combination of his account and Munn's empirical reading grounds Mechanism 3.

3.5 Contextual integrity and the integration layer

With Nissenbaum (2004) as a companion, Nissenbaum (2010) presents contextual integrity as a privacy theory based on the appropriateness of information flows in a social environment. On this account, privacy concerns whether information moved in step with the norms of the context where it was first shared and not whether it stayed secret. Information that is suitable in a doctor's office might be inappropriate in an insurance-pricing model, even if no rule has been broken. In this thesis, contextual integrity works at the integration layer. Palantir Foundry combines heterogeneous data such as police records, immigration data, financial records, and social-media metadata into a single homogeneous platform. While this process simplifies data visualisation, it equally overrides the contextual norms each source was collected under. This is the conceptual underpinning for Mechanism 1: more data does not necessarily mean better results, for the integration that creates more data also dissolves the boundaries that confer context on the data.

3.6 Classification as a cross-cutting framework

Classification runs throughout all four theories. For Bowker and Star (1999), classification systems are moral and political infrastructures, and the better they work, the less visible they become. Bowker (2005) pushes this into the database itself. By fixing what can later be recalled or contested, a database sets the limits of "potential memory". These works provide the vocabulary which the analysis uses to describe what each categorical decision at each layer means to the system. This cross-cutting framework, in combination with the four theories, is what makes it possible to read the architecture as a whole across its six layers and four mechanisms, rather than each layer by itself. The methodology for operationalising this reading is outlined in Chapter 4.

4. Methodology

This chapter lays down the manner of conducting the architectural reading in practice. The research question is causal-architectural rather than performance-evaluative, meaning that it does not ask how well an algorithm performs in a technical sense but how the layered decisions surrounding it produce the consequences observers later name as ethical issues. The method must therefore trace architecture, not measure outputs. The chapter introduces the qualitative, design-driven approach, defends it against other approaches, and defines the three components used to operationalise the framework from Chapter 3: case-study analysis, conceptual system mapping, and design probe.

4.1 Research approach and justification

The thesis takes a qualitative and interpretive approach, based on the case-study method and design probe. Yin (2018) provides the standard reference for case-study design. The thesis draws especially on his logic of single-case selection and on analytic generalisation, the appropriate inferential move when a case is theoretically rather than statistically representative. Yin's positivist stance is not compatible with the STS framing adopted here and is thus combined with Flyvbjerg (2006), who argues that the single case does not preclude generalisation. Flyvbjerg's defence of context-dependent knowledge as a legitimate research output and his insistence on the power of example underpin the inferential register the thesis claims.

4.2 Why this approach and not others

A how-question requires a method that is based on architecture. Quantitative analysis of system performance would measure outputs without explaining why they emerge from those architectures. Ethnography of users would provide access to lived experience, but not to the design logic that gave rise to the system that the users experience. Interview-based research with platform engineers or institutional clients was considered but rejected. Palantir's commercial confidentiality and security clearances required for many deployments make access infeasible at the BA level. In addition, a small sample size would not support the architectural claims of the thesis. The case study and probe together avoid these restrictions and allow the architecture itself to be examined.

4.3 Case study analysis

The empirical anchor is Palantir Foundry, examined as a single case at each of the six architectural layers. The materials are publicly available, including peer-reviewed studies (Iliadis & Acker, 2022; Brayne, 2017; Knight & Gekker, 2020; Munn, 2017; Ulbricht & Egbert, 2024), documents written by Palantir, and journalistic and policy investigations of specific deployments. Iliadis and Acker (2022) set the precedent for this thesis. They reverse-engineered Palantir's stack from 155 patents and primary documents. Their work is the central peer-reviewed source demonstrating that Palantir can be studied rigorously through public materials, without internal access.

The reading is pattern extraction rather than summary: each source is read for what it reveals about how Foundry operationalises each layer, and where the four mechanisms surface. The Palantir-authored corpus is read critically across registers. The Privacy and Governance Whitepaper (Palantir Technologies, 2024) speaks in the governance register, addressed to regulators and customers anxious about civil liberties; two AIP product pages (Palantir Technologies, n.d.-a, n.d.-b) speak in the commercial register, addressed to procurement officers and compliance managers. All three conclude with disclaimers that their screenshots and examples are notional rather than operational, and that acknowledgement is the warrant for the critical reading. The corpus is evidence of how Palantir frames its architecture for different audiences, not of what the architecture does in deployment.

4.4 Conceptual system mapping

The second component is a conceptual system mapping that traces the six architectural layers (idea, collection, integration, storage, processing, and visualisation) as a single sequential structure. The mapping is interpretive rather than technical. It identifies decision points and shows how decisions at each layer cascade into the next, but it does not diagram engineering implementation. It is based on Star's (1999) ethnography of infrastructure, which examines technical systems as sociotechnical objects whose properties become visible at points of friction and breakdown, and on Latour's (1992) following of associations between human and non-human actants. The mapping appears in Chapter 5 as a single figure that structures the layered argument and offers a visual reference for the four mechanisms.

4.5 Design probe

The third component is a design probe, the Noor System, a functional footfall-tracking system developed for this thesis. It is not a secondary case study but an instrument. It reproduces the six-layer architecture attributed to Foundry at a scale where every architectural decision can be examined and is methodologically demonstrative rather than evaluative. The build involves all six layers: a Raspberry Pi 5 and a Raspberry Pi AI Camera performing person detection on the sensor itself, with no footage being captured (collection); a self-hosted Headscale mesh returning a daily JSON log to a main server (integration); a deterministic packaging routine that writes the daily log into a NAS-hosted vector database (storage); a local Mistral 7B model, run via Ollama, as a retrieval-augmented generation (RAG) pipeline (processing); a dashboard and chatbot (visualisation). The complete specification is provided in Appendix A and the source code in Appendix B.

The point is that an architectural decision produces the same kind of effect at any scale. Following this logic, the thesis aims to demonstrate the larger architecture of Foundry, by building and examining a smaller one that works on the same principles. Using Noor as a catalyst works because every decision step is documented. Foundry, on the other hand, hides its decisions behind commercial confidentiality and security clearance. Noor is not Foundry, and the thesis never claims it is.

The ethical consequences of a data architecture are weakly emergent (Bedau, 1997). They are fully grounded in the architecture but cannot be derived from its specifications in advance. No reasoning about Foundry's design documents can establish what that design produces. A weakly emergent property becomes visible only when the system is assembled and run. The case study and the conceptual mapping describe Foundry's architecture but cannot instantiate it. The Noor System can instantiate the architecture, because it is a small architecture actually built and run, where consequences observable only in operation become traceable because every decision is documented. The methodological lineage of the probe is the cultural probe of Gaver, Dunne, and Pacenti (1999), which is presented as an interpretive design tool rather than a data-collection instrument. The term design probe is adopted from Mattelmäki (2006), who developed it from Gaver et al.'s cultural probes. It is used here in an adapted sense, denoting a built architectural artefact rather than the participant-facing self-documentation packages the term conventionally describes. Boehner et al. (2007) provide the critical companion that holds the probe to an illustrative and not evaluative standing.

The Noor System is read through its real, documented architectural choices. Every layer was built with specific decisions: no video footage collected; daily synchronisation chosen over continuous streaming, with a self-hosted Headscale mesh preferred to a commercial alternative; detection confidence scores not retained at storage; a local model constrained by curated domain knowledge at processing; and deliberate framing choices at visualisation. These are combined in Appendix A as a single decision register. Each illustrates one or more of the four mechanisms detailed in Chapter 5 and is treated as an architectural commitment, not a parameter setting. Outputs are compared qualitatively across decisions rather than benchmarked numerically.

4.6 Combining the three components

The case study grounds the argument empirically in documented deployments: the LAPD's point system (Brayne, 2017), ICE's twenty-one-database ICM (Knight & Gekker, 2020), and Hessen DATA (Ulbricht & Egbert, 2024). The conceptual mapping makes architecture visible by naming the layers and the decision points. The probe demonstrates the mechanisms in miniature, at a scale where every decision is examinable. The case study provides empirical weight, the mapping analytical structure, and the probe demonstrative isolation. Each compensates for the others' limits.

4.7 How conclusions are drawn

Conclusions are developed by repeated comparison of the case-study findings to the probe demonstrations, organised by the four mechanisms and six layers. The findings are recurring patterns: points where data integration produces consequences that map onto one or more mechanisms. The argument is qualitative and interpretive. It does not produce statistical findings in the conventional sense. The contribution is conceptual: a framework for reading architectural decisions as ethical decisions, supported by case work, and demonstrated in miniature by the probe. Flyvbjerg (2006) describes this as analytic generalisation, which rests on the power of example and on context-dependent knowledge as a legitimate research output.

5. Analysis and Discussion

5.1 Approach to the analysis

The analysis is structured by mechanism, not by layer. Each of the four mechanisms operates across one or more of the six architectural layers: idea, collection, integration, storage, processing, visualisation. Reading mechanism-by-mechanism keeps the analytical argument central, avoiding the redundancy of a layer-by-layer walkthrough. Figure 1 shows the six layers as a sequence and indicates where each mechanism is most acute; it is the conceptual map referred to throughout the chapter. The evidentiary order is the same throughout the chapter. The case-study corpus carries the empirical claim. The probe shows the same dynamic at a scale where every decision is documented: it illustrates, but it does not prove (Boehner et al., 2007). Each mechanism is presented in a similar way: a theoretical anchor establishes the conceptual claim, evidence from the Palantir case-study corpus shows where it surfaces in deployed architecture, and a demonstration through the Noor System isolates the dynamic at an examinable scale. The Palantir-authored corpus is read across registers throughout: where the governance whitepaper acknowledges a tension, the commercial AIP pages tend to erase it, and the contrast is itself evidence. The chapter closes by drawing the four mechanisms together as a single architectural argument.

Conceptual map of the data architecture as six layers of decision: idea, collection, integration, storage, processing, visualisation.
Figure 1. The data architecture as six layers of decision.

5.2 More data does not equal better outcomes

5.2.1 Integration and contextual collapse. The first mechanism operates most acutely at the integration layer. Nissenbaum's (2010) contextual integrity is the theoretical anchor: privacy is based on the extent to which information flows in line with the contextual norms under which it was originally shared. Integration platforms violate this by definition, combining sources gathered according to different norms (police records, financial records, social-media metadata, immigration data) into a unified environment in which the original constraints no longer apply. Iliadis and Acker (2022) document the collapse architecturally in Foundry's three-stage data funnel and customisable per-client ontology, which bring heterogeneous sources into a single classificatory system. Knight and Gekker (2020) describe its institutional form in the Investigative Case Management platform, which knits twenty-one databases together for U.S. Immigration and Customs Enforcement. Once sources are merged, the distinctions that gave each record its meaning are lost. Integration dissolves the context in which the data was produced. The next layer therefore receives data without the norms that made it interpretable (Nissenbaum, 2010).

5.2.2 False alignment through data richness. The sharper question is not how much data, but who it falls on. More data does not arrive evenly. It pools on whoever is already watched, so a bigger dataset confirms the existing asymmetry. Brayne (2017) finds the LAPD's Palantir-driven point system "is path dependent"; it generates a feedback loop by which FIs [field interviews] are both causes and consequences of high point values. An individual having a high point value is predictive of future police contact, and that police contact further increases the individual's point value (p. 987). The score is numeric. Points accrue with each police contact, and "quantified risk" becomes a running total that climbs because it has been climbing. People who already appear often in police data attract more attention, generate more data, justify more scrutiny. The loop starts loaded. Browne (2015) showed the inputs were already sorted along racialised lines, so the feedback compounds an asymmetry that was there before the system ran.

Palantir markets exactly this dynamic. AIP for Procurement Fraud Detection describes user feedback that "ultimately powers a feedback loop" (Palantir Technologies, n.d.-b) that refines the system's alerting rules over time. What scholars identify as a structural problem, Palantir markets as a feature: the system reinforces whatever dismissal patterns users exhibit, including unconscious bias or time pressure. Stahl's (2025) "view from nowhere" critique lands here: the feedback loop does not deliver objectivity. It systematises the patterns users already have, then presents the result as analysis and hides that fact.

5.2.3 Probe demonstration: collection and integration choices in the Noor System. Mechanism 1, aided by the Noor System, makes this visible at a scale we can examine. After the store closes, it transmits the event logs once, each day, meaning that any later analysis and its fineness are affected by the single integration choice. Therefore, the system does not represent the behaviour, e.g. peak hour, but rather produces a prediction based on the data, bucketed in days. If the transmission fails, the system either reconstructs the data from local logs or writes it off as missing. This choice does affect the regularity of the dataset.

Before any model, the packaging routine runs. It takes the day's footfall counts raw, right after the day's end. Since large language models typically work more efficiently when handed packaged data, likely JSON files that trim a document to its core, this routine is a sensible step. It is a script that writes the footfall counts next to the weather summary and marketing entry as a flat record (Appendix A). All three entries have different points of origin: the marketing entry is a plan, while the weather data is pulled from an external website, and the footfall is measured in the store. As soon as they are compacted into a single record, there is no trace of their origin. Precisely that loss of provenance is the contextual collapse. The processing layer interprets the packaged data but cannot recognise the differences. Any explanation it provides is based on a merge rule created by a person, not the actual behaviour. The issue is with the architectural choice, not the AI algorithm. Since the weather and marketing files are optional, two days can create different records even with the same footfall, and nothing later signals the difference. Following Boehner et al. (2007), the probe's purpose is to show these design choices, not to measure them. It highlights the exact decision point of contextual collapse that §5.2.1 located in Foundry's funnel, because every field is documented. Since the merge happens before the model, the explanation the operator reads back cannot be checked against the original contexts or disproven by them.

5.3 AI as amplifier, not origin

5.3.1 The locus of critique upstream. The second mechanism shifts the locus of critique upstream: AI magnifies the problems of data integration but does not produce them. For Pasquale (2015), opacity is not a flaw in the algorithm but a feature of the integrated system, produced by the architecture beneath the model. Improving the algorithm cannot fix what the architecture has already shaped.

This is the architectural reasoning Perrow (1984/1999) made canonical. Normal Accident Theory (as set out in §2.5) attributes system accidents to interactive complexity and tight coupling, not single components. Integration platforms maximise both, multiplying the connections between data sources and leaving the layers tightly coupled. The AI sits inside that arrangement as one component among many. It can amplify a consequence while not being the origin of one. In a system like this, origins are distributed across the architecture, not lodged in any single part. "Amplifier, not origin" is Normal Accident Theory carried into the ethical register.

The clearest empirical case is supplied by Knight and Gekker (2020). The harm in ICM comes from no single algorithm. The wiring of heterogeneous sources into one environment was established as contextual collapse in §5.2.1; the point here is different. Integration creates a capacity that no component possessed. Of the twenty-one interfacing objects Knight and Gekker identify, fourteen feed ICM information and seven are directly queryable, so that "a user can query and automatically import or access information from an external database without ever leaving the ICM interface" (p. 237). That cross-source query is the emergent capability; it belongs to the assembled regime, not to any single database and not to the AI on top of it, which is why improving the model cannot reach it. The same logic sits behind Foundry's data funnel and the Hessen DATA deployment, where the obstacles to regulation are architectural, not algorithmic (Iliadis & Acker, 2022; Ulbricht & Egbert, 2024). Cobbe et al. (2023) name the consequence, the "accountability horizon": "the point beyond which an actor cannot 'see', which depends on the actor and the chain" (p. 1193). Palantir's marketing confirms it inadvertently: AIP for Procurement Fraud Detection describes "AI-driven automated processes" that "take automatic action on high-risk alerts" (Palantir Technologies, n.d.-b), selling AI as decision-making on the integration layer's outputs, not as recommendation.

5.3.2 Network logic and distributed agency. Actor-Network Theory names this displacement of agency. Latour (1992) and Law (1992) argue that agency is distributed across networks of human and non-human actors. Pin responsibility on any one part, and you misrepresent how the network produces its effects. In Foundry, the AI is just one node. Around it sit data structures, ontologies, integration pipelines, forward-deployed engineers, institutional clients, and regulatory environments, and the algorithm acts only because the network has been arranged to let it.

Section 3.3 established that Foundry performs Callon's (1986) three moves; the consequence is what each one forecloses. First, translation is not the neutral merging of mixed sources. When Foundry integrates different data sources, the common semantic layer rewrites them into its own format and categories, effectively making every source speak in Foundry's terms. This is what Callon (1986, p. 214) meant by "at the end of the process, if it is successful, only voices speaking in unison will be heard". Second, enrolment then removes exit. Once analysts, datasets, and downstream applications are locked into their allocated roles, the analytical layer cannot be judged in isolation from the integration layer that feeds it. This forecloses exit, eliminating the ability to extract one piece and examine it individually. Third, the obligatory passage point removes the last option of routing critique around the architecture. Foundry becomes the single chokepoint that every institutional decision passes through, so no assessment of the model can avoid passing through the environment. Stacking the three, the result is that asking "what does the AI do?" is the wrong question, because the locus of critique is the architecture, not the model sitting on top. Yet the unison holds only while it is maintained: the alliance "can be contested at any moment. Translation becomes treason" (Callon, 1986, p. 210), which is why an architecture assembled one way can be assembled differently.

5.3.3 Probe demonstration: the Noor System's processing chain as amplifier. The Noor System's processing chain shows where amplification enters; it is in what the model does with that input: a low-confidence detection comes back out as a confident, actionable recommendation. Mistral 7B builds its recommendations based on the packaged input described in §5.2.3, so whatever was committed upstream arrives already built in. Then it amplifies what the model does with that input.

The model is also parameterised. A system prompt controls tone, length, and the recommendation types it can make. The temperature setting trades coherence and creativity. The operator sees none of this, but the same data and prompt produce different recommendations under different settings. Pasquale's (2015) observation that opacity is structural and not incidental also applies here. While §5.3.1 could only claim this for Foundry, the probe can identify the specific settings. Improving the model does not change that structural opacity. The probe reveals what Foundry, once deployed, keeps hidden under commercial confidentiality.

5.4 Classification and definition power

5.4.1 Categories as decisions about what counts. Mechanism 1 described the meaning integration removes when contexts collide. The third describes the opposite movement: the meaning the system creates and projects, through the categories it imposes and the way it renders them to the operator. The agent this time is not integration. It is the ontology, the categorical structure of the platform. The mechanism runs across the back half of the pipeline. The ontology defines, the visualisation renders, and the operator perceives and acts on what reaches them. As established in §3.6, classification is moral and political infrastructure; the stakes at this layer follow from what that infrastructure does. A category, Bowker and Star (1999) argue, "valorizes some point of view and silences another. This is not inherently a bad thing - indeed it is inescapable. But it is an ethical choice, and as such it is dangerous - not bad, but dangerous" (pp. 5–6). The force of that claim is that there is no neutral ontology to fall back to. Because every category is already an ethical choice, the response cannot be to audit the bias away but only to choose the category differently, which is an architectural intervention, not a technical one. This is why the Foundry ontology is not a neutral data model but, in Bowker and Star's terms, "frozen organizational and policy discourse" (p. 135): the categories are frozen client policy, which is what makes ICE's world differ from the LAPD's and determines what the analytical layer can later interpret. This is reinforced by Gitelman (2013): "raw data" is always already cooked, so no view that the platform offers is unstructured by the categorical commitments built into its ontology.

The architecture performs a further move. The integrated picture the operator finally sees is not more complex than the data beneath it but simpler. Organisation discards detail to produce a legible output (Holland, 1998). That manufactured simplicity is itself the danger, because a clean category or single ranked list is easy to act upon precisely because the contextual complexity that would let an operator question it has been discarded.

The dynamic ontology is the focal point of this mechanism for Foundry. As described by Iliadis and Acker (2022), it is a per-client classification system that can be customised: the categorical structure of the world according to ICE is not that of the LAPD or a private bank, and each deployment's ontological commitment determines what the analytical layer can later interpret. Stahl's (2025) reification argument applies directly: categories become things rather than tools, and what the platform outputs is an expression of the ontology's commitments. Verbeek's (2011) mediation theory pins down where the moral weight sits. When the dynamic ontology labels a person as "subject of interest" or a transaction as "high-risk", that label carries moral significance even when the data does not support the inference, because it shapes what the operator can do next. Palantir markets this capacity as speed: AIP for Export Control Compliance offers "AI-suggested Classifications" that enhance "classification efficiency" (Palantir Technologies, n.d.-a). This is a morally loaded act recast as a speed metric, the categorical weight erased.

5.4.2 Probe demonstration: classification at every layer of the Noor System. The Noor System exhibits classification power at every layer, and at each the category is set by someone other than the operator who acts on it. At collection, the Raspberry Pi AI Camera's pre-trained model fixes the categories: what counts as a person, where a passer-by ends and an entrant begins. Those rules are the training dataset's definitions, authored by people unconnected to the store, and the operator looking at the day's footfall has no access to them. The entrance itself is a line drawn on the camera frame, so "being in the store" is precisely "having crossed that line"; drawn differently, it would give different counts from the same street. At storage, the schema does not retain the classifier's confidence scores. This is not an omission but the schema author pre-deciding which questions the operator may ask: the system can report entry counts but never classifier confidence. At processing, the value judgement is explicit in the code itself. The analyzer.py thresholds that sort a day as "Excellent", "Healthy", or "Below ideal range" are, in the file's own words, "an explicit architectural decision, not a neutral measurement: drawing the line at 12 / 8 / 5 is a value judgement." The labels are prescriptive, not descriptive: a very low rate returns "exterior improvements needed", so the category does not just sort the day, it tells the operator what to do. At visualisation, the same numbers become a peak-hour bar, a heat map, or a confidence interval, each privileging an interpretation while the computation stays constant. Same data, different stories, none chosen by the person reading them.

Munn (2017) argues that Gotham's interface prefigures suspicious patterns; the Noor System's dashboard prefigures peak-hour reasoning, but the numbers are the same as in any other rendering. Classification and visualisation do not render an independent reality. They are part of the categorisation that produces the system's outputs, and treating them as neutral hides where the system's moral commitments live.

5.5 Friction as architectural intervention

5.5.1 The case for deliberate constraint. The fourth mechanism is the constructive turn. The first three are diagnostic; the fourth asks what follows. If the problem is architectural, the intervention has to be architectural too, and that claim can be grounded rather than asserted. Lessig (2006) holds that anyone choosing how to regulate behaviour "must select from among at least four modalities. Rules, norms, prices, or architecture," and that "selecting architecture as a regulator will often make the most sense" (p. 94). Transposed from his domain of internet code to data-integration architecture, the point is decisive: of the four regulators, the consequences this thesis has diagnosed are produced at the architecture layer, so only architecture reaches them. Law, norms, and the market all act on the platform from outside; architecture is the platform. Cohen (2019) makes the bridge concrete: platforms operate as "active legal entrepreneurs" (p. 48), and, this thesis adds, as architectural ones, since Palantir already exercises this kind of power. This follows from §5.2.1: integration violates contextual integrity, so constraining it by purpose corrects it. And as §5.3.1 established, accountability is a property of the supply chain, so friction belongs wherever the chain crosses contexts, actors, and jurisdictions, not only at the algorithm or the interface.

A critic could object that the fourth mechanism smuggles a value judgement of 'friction is good' into a so-far causal argument. Therefore, the sudden move from diagnosis to recommendation needs to be addressed deliberately. First, since the first three mechanisms locate the consequences in architectural decisions, any intervention must be equally located in the same decisions. Second, the claim regarding friction concerns capability, not value. §5.5.2 shows that friction is not impossible but a capability the vendor possesses and withholds. The thesis does not claim friction is costless or sufficient (§5.5.3 prices its trade-offs), and Chapter 7 presents it as a precondition for regulation, not a substitute.

5.5.2 What architectural intervention looks like in practice. Architectural intervention takes different forms at different layers. At integration, friction means purpose-binding: refusing to merge sources from incompatible contexts without explicit justification. At processing, it means justification requirements and audit logging; that is, not producing confident outputs based on low-evidence inputs and recording the basis of any inference. At visualisation, it means refusing to render confidence as visual certainty: declining to present low-evidence outputs in the same register as well-supported ones.

Palantir's presentation makes clear that the firm is aware of these differences, but that it does not use them consistently. In the governance register, the Privacy and Governance Whitepaper (Palantir Technologies, 2024) acknowledges the tension between operational efficiency and privacy and names mechanisms such as Checkpoints, Sensitive Data Scanner, and Data Lineage. In effect a friction-aware document. In the commercial register, friction is eliminated as a value: the AIP pages market classification efficiency, feedback loops that refine alerting rules, and AI-driven processes that automatically act on high-risk alerts (Palantir Technologies, n.d.-a, n.d.-b). The same company describes the same architecture in different vocabularies for different audiences: friction acknowledged as a design value in one register, treated as the obstacle the product removes in the other. The contrast itself is rhetorical evidence. The governance documents describe friction mechanisms in detail, so friction-aware architecture is clearly within the platform's technical reach, a capability that the vendor withholds depending on the audience. This defeats in advance the objection that integration platforms cannot be built with friction. Palantir's governance posture is audience-tailored, constructed when defending its architecture, dismantled when selling it. The constructive act is not about adding friction where none exists, because it already exists in the governance register. It is about extending friction consistently into the layers where commercial logic has erased it.

5.5.3 Probe demonstration: friction commitments in the Noor System. The Noor System's design builds in intentional friction at multiple layers. The most consequential is at collection because the system has no video-recording capability at all. The Raspberry Pi AI Camera runs detection on-device and outputs only structured event logs. Nothing else survives: no stored footage, no extracted facial features, and no way to re-classify a detection by going back to source video, as there is none to begin with. The choice is irreversible. The system cannot be turned into a tool to identify individuals or feed a face-recognition pipeline, because the footage never existed. This is architectural intervention: the constraint is built into the structure, not promised in later applied policy. A second commitment operates at integration: the system uses Headscale, a self-hosted mesh, instead of the commercial Tailscale. Tailscale would be operationally simpler but transmits metadata to its vendor. Choosing Headscale is therefore a friction commitment in Lessig's (2006) sense. A third operates at processing: the local Mistral model never transmits inference data to a third-party provider, unlike more powerful models hosted in the cloud, which would expose the architecture's queries to the model vendor.

These commitments come with a price. The no-footage decision means classification errors cannot be retroactively audited; the self-hosted mesh imposes setup complexity; the local model limits the system to seven billion parameters rather than a frontier system. Each is a trade-off, and the trade-offs are visible in the Noor System as they would not be in a closed commercial deployment. The probe demonstrates two things. First, architectural friction is implementable; this is the constructive move Chapter 7 defends. Second, its costs are bounded and tractable, at least at this scale.

5.6 What the analysis demonstrates

The analysis has now shown what the introduction could only assert: the consequence is weakly emergent in the sense set out at §1.3, grounded in the architecture yet not derivable from its specifications in advance. The result is methodological: because the consequence appears only in the interaction of the layers, no component-level audit can locate it, only a reading of the architecture. The point is the arrow between the mechanisms, not the list of them: a property that only the architectural reading catches because only the interaction produces it. Section 6.3 takes this to the case where the emergence becomes literal, and Chapter 7 to the redesign it requires.

6. Limitations and Ethical Considerations

The argument is limited by the materials available, the scope of the case, the ethics of working with surveillance technologies at a distance, and the perspective of the piece. This chapter defines those limits, and the credibility of the argument lies in its statement of what it does not claim, as well as in what it does claim.

6.1 Methodological limitations

The thesis relies on publicly available material about Palantir. Internal architectural documentation is not accessible to a BA student researcher: engineering specifications, deployment-specific configuration files, contract terms, operational logs. Two bodies of work partly offset this. Iliadis and Acker's (2022) reverse-engineering of 155 patents is the most thorough public account of Palantir's stack to date, and the journalistic and policy investigations cited throughout add further detail. Neither fully substitutes for internal access. Claims about how Foundry behaves in a particular deployment are inferred from public traces, not observed directly.

Although the design probe serves as an illustration to examine the mechanisms (§5.1, Boehner et al., 2007), that very illustrative status is itself the limitation. The researcher developed the four mechanisms from the case study and then built the probe to demonstrate them. Therefore, the probe cannot disconfirm what it was designed to illustrate: the circular problem. The claim could still be proven wrong. If a deployed system showed ethical consequences clearly traceable to a single component, and removable by fixing that component alone, the emergent-architecture claim would weaken. The documented record reviewed in Chapter 5 contains no such case. The Noor System's documented architectural decisions show that the shared premise of integration-as-default produces specific kinds of consequence; they do not establish that Foundry produces those consequences with the same magnitude or in the same form. The qualitative, interpretive method does not produce statistical findings. The contribution is what Flyvbjerg (2006) calls analytic generalisation (§4.1, §4.7): it generalises from a well-understood case by reasoning, not by the statistical prediction a quantitative study would offer.

6.2 Scope limitations

The case is Palantir Foundry specifically, with supporting reference to Palantir Gotham (Munn, 2017) and Investigative Case Management (Knight & Gekker, 2020). Other large-scale integration systems (Clearview AI, NEC's biometric platforms, national-security data platforms outside the United States) are not examined. The architectural reading is intended to transfer to such cases, but that transfer has not been shown. Its conclusions should be read as anchored in the Palantir corpus, not as general statements about every integration platform. The legal and regulatory aspects of Palantir's implementation are touched on through Ulbricht and Egbert (2024) and Cohen (2019). However, they are not examined in detail because the thesis is sociotechnical, not doctrinal, and a serious legal analysis would require a different methodology. The argument applies most directly to high-stakes governance and security contexts; its transferability to healthcare, education, employment, or financial services is implied by the architectural framing but not demonstrated.

6.3 Ethical considerations and the scaling argument

The thesis involves no human subjects. The Noor System operates on non-identifiable data; no customer information or institutional records are processed. The ethical risks of the project as a research artefact are therefore largely intellectual: the risk of oversimplifying Palantir's complexity into a six-layer schema, and the risk of producing critique that, however unintentionally, normalises surveillance systems by treating them as redesignable rather than objects of more fundamental contestation. The first is addressed by providing evidence from a case study and presenting the probe as illustrative, while the second is addressed reflexively in Section 6.4.

The more consequential ethical consideration emerges not from the thesis as a project but from what the Noor System reveals about its subject matter. The Noor System was intentionally designed to be ethical: all the architectural decisions listed in Appendix A (no footage at collection, local Mistral service at processing, Headscale mesh instead of commercial alternatives, no personally identifiable data in the schema) were made to limit what it could become later. It is, by design, the most ethical version of itself. And yet the same architecture replicated across fifty stores in one city would become something no single one of those commitments could prevent. A system that records no faces and stores no identifying information at any single point becomes, deployed across an urban environment, a comprehensive map of pedestrian behaviour: who passes which doors, when, in what numbers, in conjunction with which weather conditions, marketing campaigns, and demographic profiles.

Every individual instance is as ethical as the original. The aggregate is not the same. This is emergence in its purest form (Holland, 1998): fifty instances are not the same as one instance fifty times. The aggregate has a property, a continuous behavioural map of an urban environment, which none of the deployments have and none of the per-instance ethical commitments can address, since it is an emergent property of the combination. Surveillance critique refers to that collection as the all-seeing eye: the privatised infrastructure of urban observation that no civic body authorised, and no individual can opt out of. The four mechanisms of Chapter 5 remain the same in the city-wide deployment as in the one-store deployment; the difference is the surface area on which they operate. This is not a speculation about a future deployment, but a generalisation in Flyvbjerg's (2006) sense, using the probe's documented commitments to make visible what the same commitments would produce at scale. The implication for the central claim is direct. Architectural intervention is the only thing that distinguishes the benign instance from the problematic aggregate, because the aggregate is what the architecture, once replicated, becomes, and no commitment made within a single deployment can prevent it.

What would prevent it is structural: purpose-binding at the integration layer, friction at the layers where deployments could combine, and regulatory limits on aggregation. The ethical question raised by this thesis is not whether the Noor System is ethical at one store, but whether the architecture it reproduces should be permitted to scale at all without the friction commitments that keep benign instances from aggregating into a problematic network. The answer the thesis defends is that it should not, and the constructive move toward architectural friction is therefore not optional but constitutive of any deployment that takes its ethical commitments seriously.

6.4 Reflexivity

The author is a Digital Society student writing from a European institutional context (Maastricht University, Faculty of Arts and Social Sciences) about a company headquartered in the United States whose deployments span multiple jurisdictions. This shapes both access and framing. Access is limited to public sources in English and German, and the analysis cannot draw on the fieldwork an institutionally embedded researcher might pursue. The framing is sociotechnical and philosophy-adjacent rather than legal, technical, or activist. This thesis does not claim a view from nowhere: its argument is offered from this specific scholarly location, and that situatedness is part of why the contribution is conceptual rather than empirically comprehensive. One boundary is worth naming because the framework itself points to it. Starting the six layers at the idea is itself a classificatory choice: the framework foregrounds architecture and brackets the prior question of what law permits, which means it valorises one point of view and silences another in exactly the sense §5.4.1 used against the platforms it studies. This thesis treats that policy landscape as a given. The constructive move of the next chapter, redesign rather than rejection, is offered as one position within a wider debate, not as a settled prescription.

7. Conclusion

7.1 Returning to the central claim

The claim established across Chapters 1 and 5 can now be treated as settled: the consequences this thesis has examined are emergent properties of the data architecture. The introduction and literature review located the gap between normative AI ethics and technical optimisation, and the analysis showed why neither reaches consequences that arise from the interaction of layers rather than from any single one. That is the gap an architectural reading fills.

7.2 What the analysis demonstrated

The four mechanisms function as one argument. Integration violates contextual integrity at the layer where heterogeneous sources are merged, and volume redistributes visibility, rather than reducing uncertainty. The algorithm above the integration cannot correct this: its agency is distributed across a mixed network whose properties act alongside its human deployers. The categorical structure integration imposes is itself consequential. Classification is a decision about what counts, and visualisation renders those decisions as the patterns operators come to see. The constructive response is friction: deliberate constraint at integration, processing, and visualisation.

Together, the mechanisms demonstrate the sense in which the consequences are emergent (Holland, 1998): they belong to no single layer but arise from the interaction of all of them, which is why a better algorithm cannot remove a consequence the architecture produces. Winner's (1980) claim that technology is not neutral grounds the reading; Bowker and Star's (1999) account of classification as moral infrastructure and Nissenbaum's (2010) contextual integrity are its sharpest layer-specific extensions.

7.3 Architectural intervention as the path forward

The constructive move is redesign, not rejection. Integration platforms exist because they answer real coordination problems, and they are deployed at a scale where pretending they will be abolished is not a more rigorous position than engaging with how they should be structured. The intervention must be architectural because the consequences are, and Lessig's (2006) argument that code is law, extended by Cohen (2019) into platform power as quasi-legislative, makes design choices governance choices.

Redesign rather than inspection is the precise form of the move. An emergent consequence cannot be lifted out of a finished system the way a fault is repaired; it can only be prevented by changing the conditions from which it grows. The architecture is a cultivated thing (its conditions are chosen), and that is exactly what makes redesign possible: what was set up deliberately can be set up differently. Friction is the constructive principle: purpose-binding at integration, justification and audit logging at processing, the refusal to render confidence as visual certainty at visualisation. Palantir's own governance vocabulary shows that this friction is already articulated when the audience demands it; the constructive move is to extend it consistently into the registers where commercial logic has erased it. Architectural friction is not a substitute for regulation but its precondition: a law mandating purpose-binding cannot be enforced if the architecture has no mechanism for purpose-binding to attach to.

This is not optional. As the scaling argument of Section 6.3 showed, friction commitments are the conditions under which an architecture can be replicated without aggregating into the surveillance apparatus this thesis has analysed. The difference between the benign single instance and the problematic aggregate lies not in per-instance design but in whether structural friction at the layers where instances combine has been built in. Architectural friction is therefore not a feature added to a working system, but the condition under which it can work without producing the consequences this thesis has diagnosed.

7.4 Directions for further research

The argument invites several extensions. Empirical study of how architectural interventions perform under real regulation, particularly the European Union's AI Act (Regulation (EU) 2024/1689), would test whether friction-by-design survives at scale, or whether, as critics will reasonably argue, it gives way to commercial pressure. One observation frames that direction: the Act regulates systems and their outputs, yet this thesis has located the consequences upstream of both. A further line of work is comparative, and the comparison cases named in §6.2 are not interchangeable. Clearview AI is a single-modality face-matching system rather than a multi-source integration platform; the non-US national-security platforms sit under different regulatory regimes. The two cases therefore let the comparison vary architecture and regulatory context separately: Clearview isolates architecture, the non-US platforms isolate regime. That is what makes it a genuine test, showing whether the four mechanisms are properties of integration as such or artifacts of Foundry's specific design. A further direction comes from turning the framework on itself. As §6.4 noted, the reading begins at the idea layer and leaves untouched the prior question of what law permits. This thesis argued, with Lessig (2006) and Cohen (2019), that architecture does the work of governance.

Following the groundwork the thesis built, the natural extension is to ask how governance shapes what gets built. This direction might be called policy-by-design. Cavoukian (2011) explains privacy-by-design that builds values into a system that already exists. It is not the same as policy-by-design that names the landscape that decides which systems come to exist at all. Closer engagement with the legal and regulatory dimensions touched on here would convert the architectural examination into a governance proposal legislators could act on. Especially at a time when the European Union acts rather reactively, introducing a new act for each new problem (GDPR, DSA, DMA, AI Act), then additional omnibus packages to reconcile them after the fact (Graux et al., 2025). It seems only logical to draft regulations that already restrict the technology's architecture while building, to prevent future issues.

This thesis can be seen as the foundation for such. It examines the middle ground of surveillance studies and AI ethics, introducing a framework of four mechanisms across six architectural layers to guide future research.

References

Bedau, M. A. (1997). Weak emergence. Philosophical Perspectives, 11, 375–399. https://doi.org/10.1111/0029-4624.31.s11.17

Bijker, W. E., Hughes, T. P., & Pinch, T. (Eds.). (1987). The social construction of technological systems: New directions in the sociology and history of technology. MIT Press.

Boehner, K., Vertesi, J., Sengers, P., & Dourish, P. (2007). How HCI interprets the probes. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems (pp. 1077–1086). ACM. https://doi.org/10.1145/1240624.1240789

Bowker, G. C. (2005). Memory practices in the sciences. MIT Press.

Bowker, G. C., & Star, S. L. (1999). Sorting things out: Classification and its consequences. MIT Press.

boyd, d., & Crawford, K. (2012). Critical questions for big data: Provocations for a cultural, technological, and scholarly phenomenon. Information, Communication & Society, 15(5), 662–679. https://doi.org/10.1080/1369118X.2012.678878

Brayne, S. (2017). Big data surveillance: The case of policing. American Sociological Review, 82(5), 977–1008. https://doi.org/10.1177/0003122417725865

Browne, S. (2015). Dark matters: On the surveillance of blackness. Duke University Press.

Callon, M. (1986). Some elements of a sociology of translation: Domestication of the scallops and the fishermen of St Brieuc Bay. In J. Law (Ed.), Power, action and belief: A new sociology of knowledge? (pp. 196–233). Routledge & Kegan Paul.

Cavoukian, A. (2011). Privacy by design: The 7 foundational principles. Information and Privacy Commissioner of Ontario.

Cobbe, J., Veale, M., & Singh, J. (2023). Understanding accountability in algorithmic supply chains. In Proceedings of the 2023 ACM Conference on Fairness, Accountability, and Transparency (FAccT '23) (pp. 1186–1197). ACM. https://doi.org/10.1145/3593013.3594073

Cohen, J. E. (2019). Between truth and power: The legal constructions of informational capitalism. Oxford University Press.

Flyvbjerg, B. (2006). Five misunderstandings about case-study research. Qualitative Inquiry, 12(2), 219–245. https://doi.org/10.1177/1077800405284363

Gandy, O. H., Jr. (1993). The panoptic sort: A political economy of personal information. Westview Press.

Gaver, B., Dunne, T., & Pacenti, E. (1999). Design: Cultural probes. Interactions, 6(1), 21–29. https://doi.org/10.1145/291224.291235

Gitelman, L. (Ed.). (2013). "Raw data" is an oxymoron. MIT Press.

Graux, H., Garstka, K., Murali, N., Cave, J., & Botterman, M. (2025). Interplay between the AI Act and the EU digital legislative framework (PE 778.575). European Parliament, Policy Department for Transformation, Innovation and Health. https://www.europarl.europa.eu/RegData/etudes/STUD/2025/778575/ECTI_STU(2025)778575_EN.pdf

Holland, J. H. (1998). Emergence: From chaos to order. Addison-Wesley.

Iliadis, A., & Acker, A. (2022). The seer and the seen: Surveying Palantir's surveillance platform. The Information Society, 38(5), 334–363. https://doi.org/10.1080/01972243.2022.2100851

Iliadis, A., & Russo, F. (2016). Critical data studies: An introduction. Big Data & Society, 3(2), 1–7. https://doi.org/10.1177/2053951716674238

Kitchin, R. (2014). Big data, new epistemologies and paradigm shifts. Big Data & Society, 1(1), 1–12. https://doi.org/10.1177/2053951714528481

Knight, E., & Gekker, A. (2020). Mapping interfacial regimes of control: Palantir's ICM in America's post-9/11 security technology infrastructures. Surveillance & Society, 18(2), 231–243. https://doi.org/10.24908/ss.v18i2.13268

Koskinen, J., Knaapi-Junnila, S., Helin, A., Rantanen, M. M., & Hyrynsalmi, S. (2023). Ethical governance model for the data economy ecosystems. Digital Policy, Regulation and Governance, 25(3), 221–235. https://doi.org/10.1108/DPRG-01-2022-0005

Latour, B. (1992). Where are the missing masses? The sociology of a few mundane artifacts. In W. E. Bijker & J. Law (Eds.), Shaping technology / building society: Studies in sociotechnical change (pp. 225–258). MIT Press.

Law, J. (1992). Notes on the theory of the actor-network: Ordering, strategy, and heterogeneity. Systems Practice, 5(4), 379–393. https://doi.org/10.1007/BF01059830

Lessig, L. (2006). Code: Version 2.0. Basic Books.

Lyon, D. (2022). Surveillance. Internet Policy Review, 11(4). https://doi.org/10.14763/2022.4.1673

Mattelmäki, T. (2006). Design probes [Doctoral dissertation, University of Art and Design Helsinki]. University of Art and Design Helsinki.

Mittelstadt, B. D., Allo, P., Taddeo, M., Wachter, S., & Floridi, L. (2016). The ethics of algorithms: Mapping the debate. Big Data & Society, 3(2), 1–21. https://doi.org/10.1177/2053951716679679

Munn, L. (2017). Seeing with software: Palantir and the regulation of life. Studies in Control Societies, 2(1). https://studiesincontrolsocieties.org/seeing-with-software/

Nissenbaum, H. (2004). Privacy as contextual integrity. Washington Law Review, 79(1), 119–157.

Nissenbaum, H. (2010). Privacy in context: Technology, policy, and the integrity of social life. Stanford University Press.

Palantir Technologies. (2024). Palantir privacy and governance whitepaper. https://www.palantir.com/assets/xrfr7uokpv1b/6pey1VnYHULqeggNbPKqP0/9f577de3e3dfb9fc031bd75dc7526517/Palantir_Privacy_and_Governance_Whitepaper__1_.pdf

Palantir Technologies. (n.d.-a). AIP for export control compliance. Retrieved 19 June 2026, from https://aip.palantir.com/workflow/ab899ed9-fead-4d99-ad88-6fdc01045bb4

Palantir Technologies. (n.d.-b). AIP for procurement fraud detection. Retrieved 19 June 2026, from https://aip.palantir.com/workflow/38a952df-f1d9-4782-887e-b3f5a5abeeb8

Pasquale, F. (2015). The black box society: The secret algorithms that control money and information. Harvard University Press.

Perrow, C. (1999). Normal accidents: Living with high-risk technologies (Updated ed.). Princeton University Press. (Original work published 1984)

Plantin, J.-C., Lagoze, C., Edwards, P. N., & Sandvig, C. (2018). Infrastructure studies meet platform studies in the age of Google and Facebook. New Media & Society, 20(1), 293–310. https://doi.org/10.1177/1461444816661553

Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union. http://data.europa.eu/eli/reg/2024/1689/oj

Singh, J., Cobbe, J., & Norval, C. (2018). Decision provenance: Harnessing data flow for accountable systems. IEEE Access, 7, 6562–6574. https://doi.org/10.1109/ACCESS.2018.2887201

Sismondo, S. (2010). An introduction to science and technology studies (2nd ed.). Wiley-Blackwell.

Stahl, B. C. (2025). The ethics of data and its governance: A discourse theoretical approach. Information, 16(6), Article 497. https://doi.org/10.3390/info16060497

Star, S. L. (1999). The ethnography of infrastructure. American Behavioral Scientist, 43(3), 377–391. https://doi.org/10.1177/00027649921955326

Ulbricht, L., & Egbert, S. (2024). In Palantir we trust? Regulation of data analysis platforms in public security. Big Data & Society, 11(3), 1–15. https://doi.org/10.1177/20539517241255108

Verbeek, P.-P. (2011). Moralizing technology: Understanding and designing the morality of things. University of Chicago Press.

Winner, L. (1980). Do artifacts have politics? Daedalus, 109(1), 121–136.

Yin, R. K. (2018). Case study research and applications: Design and methods (6th ed.). SAGE Publications.

Zuboff, S. (2019). The age of surveillance capitalism: The fight for a human future at the new frontier of power. PublicAffairs.


Noah Najafi — bachelor thesis, B.A. Digital Society, Faculty of Arts and Social Sciences, Maastricht University. First published June 2026.