Stage 4 summary. Information Systems Architecture
Stage 4 is TOGAF Phase C worked through end to end. It describes what information the enterprise holds, who governs it, and which applications carry responsibility for it, and it develops the data view and the application view as one paired effort rather than two parallel projects. The pairing is deliberate, because authority, integration, and publication choices are coupled, and a clean answer in one view is worthless if it contradicts the other.
The thread running through every module is that information decisions are governance decisions before they are tooling decisions. Authority, meaning, and traceability must be made explicit, and the stage's working test is blunt. A governance board should be able to trace any published figure back to its authoritative source through every layer of the Phase C pack. If a link in that chain is missing, the pack is a collection of documents rather than an architecture.
The sections follow the stage's teaching order, so you can read straight through to rebuild the stage in your head, or jump to the concept you need. Each section links back to its module for the full treatment.
What you carry out of this stage
- Explain why Phase C develops data and application architecture together and what separating them produces at the boundaries
- Build an information map whose domains are units of business information responsibility, each with a recorded data architecture pattern
- Replace single source of truth slogans with a five-layer authority map covering entity, attribute, authority, stewardship, and publication
- Settle the six customer MDM decisions before any platform choice, adapted to the connection-based reality of utility customers
- Treat asset, network, planning, telemetry, and publication information as enterprise domains, with CIM alignment decisions recorded with rationale
- Attach the six minimum metadata concerns to every published output and design analytics as a four-layer decision-support chain
- Define ABBs before SBBs and record every integration decision with its interoperability level, coupling choice, and failure consequences
- Trace a published figure end to end through the Phase C pack, and govern analytics and AI models as architecture elements
Phase C develops data and applications as one domain
Phase C is the ADM phase where information architecture and application architecture are developed as one coordinated domain. C220 Part 1 gives it two canonical objectives. Develop the target Information Systems (Data and Application) Architecture in a way that enables the Business Architecture and the Architecture Vision, and identify candidate Architecture Roadmap components from the gaps between baseline and target. The standard works through those objectives twice, once for data and once for applications, as two parallel sequences that share the same seven-step shape, from target architecture through ABB identification and development, stakeholder review, gap analysis, and principle alignment to formal documentation.
The pairing exists because the decisions are mutually dependent. Information authority decisions constrain application boundaries, and application responsibilities determine how information governance is enforced. Run the two as separate workstreams and each side produces a strong target-state document whose authority assignments contradict the other, a conflict nobody discovers until a governance review asks who is authoritative after go-live. Phase C also never starts from a blank page. It inherits capabilities and value stages, business gaps, stakeholder concerns, architecture principles, and business scenarios from the Stage 3 business architecture, and weak inputs should be flagged before work proceeds.
Opening the phase with product or platform evaluation starts too low in the stack, because products belong after the enterprise has named its information responsibilities and application boundaries. Five Series Guides turn the generic method into a working toolkit for the stage, G190 for information mapping, G21B for customer master data, G234 for metadata, G238 for analytics, and G248 for selecting building blocks, and the first-pass pack they support should be readable by non-technical stakeholders.
The two parallel tracks of TOGAF Phase C and their joint gap output
Phase C takes the Phase B handover and runs data architecture and application architecture as two equal workstreams at once, then merges them into one integrated gap analysis, so Phase D gets a single signed-off architecture, not a data gap and an application gap settled apart.
Information domains map meaning, not databases
An information map groups the enterprise's information into business-relevant domains, each with named ownership, stewardship, and consumer relationships. It is broader than any schema and more structured than a loose glossary, and each domain is a unit of business information responsibility that survives even when the underlying systems are replaced. The London map names customer, connection, planning, asset, network, telemetry, and publication domains, groupings a business sponsor can recognise without a data-dictionary lookup.
The practical technique has five steps. Start from the business capabilities and value stages identified in Stage 3, name each domain in enterprise language, identify the major consumers and their decision needs, record the key relationships between domains, and note where the current estate aligns or conflicts with the domain structure. The failure mode is the relabelled inventory, a map whose domain names contain system names and whose boundaries mirror the current applications, which inherits the accidents of the existing estate and ages with it.
This course uses four data architecture patterns to reason about how each domain is implemented, centralised, federated, replicated, and hybrid. They are a practical industry vocabulary rather than a taxonomy from any single TOGAF document, and the pattern choice for each domain carries different governance, latency, consistency, and cost consequences, so it should be recorded and justified in the Phase C pack.
Information domains group data by business responsibility
A hospital's information splits into five domains named by what the business is responsible for, not by the databases that store it, so a domain keeps its name and its owner when the system of record behind it is replaced.
Customer master data is decided before any platform
Customer master data management is a set of architecture decisions, not a box to be bought and switched on. G21B identifies six decisions that must be settled before platform selection. Entity definition, shared-attribute identification, authority assignment, matching rules, stewardship, and the decision use case the whole effort serves. Programmes that promise one master system as a one-time clean-up skip exactly the identity logic and authority boundaries that make the result trustworthy.
Golden-record language earns its place only when the underlying choices are explicit. A golden record becomes an architecture artefact when it specifies what it contains, how conflicts between sources are resolved, and who is responsible for it. Tied to specific high-value decisions and authority problems, master data thinking controls customer information without flattening the genuine differences across the enterprise.
Utility customer data needs the retail model adapted before any platform implements it. Relationships are connection-based, defined by a physical connection point rather than a purchase, and the same entity can be a connection applicant, a service consumer, a metering-point holder, and a complaints correspondent across a lifecycle spanning decades, with regulatory reporting obligations attached. The working test of the architecture is trust. Shadow copies of customer data are a symptom of low trust, and if the architecture is working the need for them should decrease.
The customer master data lifecycle that closes back on itself
A customer record is created, matched, merged, governed and retired, and the retired record loops back as context for the next intake, so a returning customer is recognised rather than created afresh and the master never silts up with duplicate dead records.
Asset and network data are enterprise domains, and CIM is a vocabulary
For an asset-heavy operator, five information groupings carry the architectural weight. Asset, network, planning and scenario, telemetry and operational, and publication information. The same data feeds planning, safety, resilience, and regulatory evidence, which makes model authority and data quality architecture concerns rather than technical housekeeping. The stage's cautionary case is a GIS and an asset register disagreeing about a cable route, with a planning estimate coming out materially wrong because no architecture governed the relationship between systems holding different aspects of the same physical reality.
The Common Information Model, standardised as IEC 61968 and IEC 61970, is a shared semantic vocabulary for power-systems information, not a mandatory schema. Alignment does not mean every system uses CIM natively. It means the enterprise knows the mapping between its internal structures and the CIM, governs that mapping at the boundaries where information crosses to external or publication formats, and records where alignment is maintained and where it is deliberately omitted, with rationale, in the Phase C pack.
The trap is letting operational-system boundaries dictate the enterprise information architecture. A system can be excellent at running a function and still be a poor organising principle for publication, analytics, and governance, so asset and network data are treated as enterprise domains in their own right.
Every asset data domain conforms to one authoritative model
For a railway operator each network-asset data domain traces from the store it lives in through the authoritative model that store must conform to, to the operational consumer that reads it, so no operational system can dictate the enterprise information architecture.
Metadata and analytics carry trust, not decoration
Metadata supplies the context that raw data cannot carry on its own, and G234 sets six minimum concerns every published or shared information asset should answer. Meaning, provenance, stewardship, lifecycle, classification and access, and quality and confidence. Technically accurate substation capacity figures with no scenario, model-run date, or authority context are practically misleading, and without metadata the enterprise quietly outsources interpretation risk to every downstream user. Retrofitting it later costs more, because the original knowledge holders have moved on, the asset count has grown, and consumers have already built their own interpretations. Publication raises the burden further, since external consumers cannot ask the data team for context, and the decisions about who may see and receive which figures sit alongside the metadata as design choices, with the paired baseline-to-target data view stating how each domain moves to the target without loss or duplication.
Analytics is a decision-support chain, and G238 structures it in four layers. Source information with its authority and quality constraints, the semantic layer of shared definitions and calculation logic, the analytical product the consumer sees, and the decision use that gives the chain its purpose. The semantic layer is the most common failure point because it is the least visible and the most easily skipped. When fourteen dashboards, or a dashboard and a regulatory submission, disagree on the same concept, the cause is almost always an ungoverned semantic layer.
Trustworthy analytics is therefore built from the foundations up. G238 describes four patterns, centralised, federated, self-service, and embedded, suited to different decision needs, and refresh frequency must match the decision cadence the product serves. A dashboard is an output of the architecture, not the architecture, and an analytical product with no named decision use is an output without a customer.
The BI chain from raw source data to a named board decision
Analytics architecture is judged by the decisions it supports, not by how a chart looks: source feeds a model, the model serves a dashboard, an analyst reads an insight, and the board decides, so a dashboard that never reaches a decision is decoration.
ABBs name the responsibility before SBBs name the product
Application architecture describes the responsibilities the enterprise needs, not the products it happens to own. An architecture building block is a logical description of a needed capability, and C220 Part 4 makes it technology-aware but product-neutral. It captures functional and non-functional requirements, defines interfaces to other ABBs, guides and constrains the selection of solution building blocks, and is reusable across contexts. A solution building block is the product-specific implementation that realises one or more ABBs and is constrained by their specifications. Naming a vendor platform as the ABB means the logical layer has been skipped and the vendor's pitch has become the de facto requirement.
The content metamodel, renamed the Enterprise Metamodel in the 10th Edition, is what stops the application view floating free. It keeps application components traceable to the data entities they create, read, update, or delete, the business functions and capabilities they support, the technology services that will host them, and the mapping from logical to physical component that ties every product choice back to an enterprise need.
The London ABB catalogue shows the discipline in use. Planning-evidence preparation, publication management, telemetry ingestion, customer-data stewardship, network-model management, and governance reporting are named as enterprise responsibilities first, so candidate products can be weighed against a stable logical requirement instead of procurement pressure standing in for architecture reasoning.
The ABB to SBB selection trace from pattern to delivery
A Solution Building Block never lands without an Architecture Building Block above it to justify it: the architect selects a pattern, the board evaluates it, and only the realised solution, the single step that carries cost and vendor lock-in, is contracted into delivery.
Integration and coupling decide how the enterprise can change
Integration is an architecture concern because coupling decisions drive failure propagation, change cost, and accountability. The stage's two cautionary numbers make the point. A synchronous, tightly coupled telemetry pipeline lost ninety minutes of sensor readings when the receiving system went offline, and an estate of seventy-two independently designed point-to-point integrations made a single system replacement touch thirty-one of them. One badly chosen pattern is harmless. Dozens create a fragile web that makes every change programme expensive.
The TOGAF Standard frames interoperability by how business processes, information, and technical services are shared, and asks architects to state the degree of information exchange each interaction needs. On top of that framing this course works with four levels, operational, syntactic, semantic, and organisational, as a teaching taxonomy rather than a list from the standard, and each integration should specify which level it requires. Semantic interoperability is where most enterprise integration problems actually live, because two systems can exchange data successfully while interpreting it differently.
Coupling choices follow the enterprise context, not technical preference, weighed by failure consequence, change frequency, latency requirement, and governance accountability. A good integration decision record captures five things. The business purpose of the connection, the interoperability level required, the authority and semantic assumptions travelling across it, the chosen pattern with the alternatives considered, and the consequences for latency, failure handling, coupling, and governance, with an explanation of why the chosen pattern is the least harmful option.
Integration coupling across synchrony and contract strength
Coupling is set by two independent choices: whether the call is synchronous or asynchronous, and whether the contract is loose or strong. Synchronous plus strong (RPC, gRPC) is the tightest binding on the board, so pick it only where the use case cannot tolerate delay.
The London walkthrough proves the pack with one traceability test
The walkthrough assembles everything the stage taught into one Phase C pack, built in a six-layer sequence where each layer answers a question the next depends on. The information-domain map, the authority map, the CIM alignment decisions, the metadata and publication rules, the application responsibilities and key interactions, and the integration and governance notes. Starting with application responsibilities before the domains are clear produces platform-driven answers, which is why the order matters.
The strongest validation is end to end. A network capacity headroom figure in the LTDS publication carries metadata naming its scenario, model-run date, steward, and confidence. From there the trace runs back through the publication ABB and its SBB mapping, the CIM alignment at the publication boundary, the governed integration path with its decision record, the authority chain for planning assumptions and inputs, and finally the source domains that fed the forecast. If a governance board can follow that whole trace, the pack has done its job. If any link is missing, the pack is incomplete.
The pack then anchors everything that follows. Stage 5 selects the technology services that host the application responsibilities defined here, Stage 6 sequences change in an order that respects the authority boundaries and integration dependencies the pack records, and Stage 7 governs implementation against it as the reference architecture.
The six hidden steps behind a published LTDS product
Between the operator's systems and the regulator's desk the LTDS clears six owned steps, mapped to CIM, validated, packaged, published and consumed, so any of the five steps the reader never sees is still a place the data can be wrong and needs its own owner and check.
Analytics and AI models are architecture elements, governed in the repository
Analytics and AI are extensions of the information architecture, not a separate discipline. A model is an application built on data, so it belongs to Phase C, and the moment its output changes a decision it carries authority like any other authoritative source. The stage's opening warning is a load-forecasting model that shaped reinforcement spending for eighteen months with no owner, no recorded lineage, and no review point. The remedy is to treat the AI system, the model plus its training data, serving code, and the decision it feeds, as one architecture element with an owner, a stated purpose, a lifecycle, a repository entry, and explicit dependencies.
Governance separates into three questions that form a chain, not a menu. Data governance asks whether the input is fit to use, traced to sources with a recorded basis. Model governance asks whether the model is sound, versioned, and evaluated across the groups it affects rather than only in aggregate. Decision governance asks whether a human can review and override the live decision, whether affected people can challenge it, and whether drift is monitored with a defined rollback. Good data with an ungoverned model, or a sound model with no decision control, still fails.
Four controls carry the load, all ordinary architecture artefacts. Data lineage, a model registry, human oversight, and a risk tier that sets how heavy the other controls must be. Regulators are moving to classify AI by risk, so the risk-based shape is worth building in, but the precise legal categories, duties, and timelines must be verified against the current text with qualified advice rather than asserted from a course. London governs its load-forecasting model through the repository, so the same record that assures the model doubles as Ofgem investment evidence.
The AI governance control chain
An AI model reaches a real decision only after five architecture controls, data lineage, model registry, an evaluated risk tier and human oversight, and each records its evidence in the repository, so an ungoverned model cannot quietly reach a decision no control could trace.
One published figure, traced all the way down
The stage closes by running its own test on London Grid Distribution, the fictional operator serving 2.3 million customers across Greater London. A capacity headroom figure published under LTDS obligations must be traceable from the consumer back to its sources. The metadata names the scenario and model-run date, the publication ABB and its platform prepared it, CIM alignment made it interpretable at the boundary, a governed integration carried it, the authority map shows planning-assumption authority sitting in the modelling tool while publication authority sits in the LTDS layer, and the domain map shows the asset, network, planning, and telemetry information that fed it. Every layer of the stage appears in that one trace.
The same discipline reaches the newest element in the estate. The AI load-forecasting model that ranks substations for reinforcement is a versioned repository entry with an owner, recorded lineage, a planner reviewing its output before money moves, and drift monitoring with a defined rollback, so the record that assures the model is the record that justifies the investment case to Ofgem. Hold the test in mind, because every stage that follows will be asked the same question under different pressure. Can you prove why what you publish deserves to be trusted?
The traps this stage warns against
Opening Phase C with product, platform, or integration-tool evaluation.
Instead: Name the information domains, the authority assignments, and the application responsibilities first. Products are weighed against those, never the other way round.
Declaring a single source of truth for a whole entity.
Instead: Different attributes, lifecycle stages, and publication views are mastered by different systems and roles, so record a five-layer authority map and bound any single-source claim by entity, attribute, consumer, and purpose.
Naming information domains or ABBs after the systems and vendor platforms that currently hold the data.
Instead: Define domains as units of business information responsibility and ABBs as enterprise responsibilities in enterprise language, so the architecture survives system change instead of locking in the existing estate.
Treating publication, metadata, or integration as add-ons bolted on after the main work.
Instead: They are upstream Phase C layers. Build the six metadata concerns, the publication rules, and the integration decision records into the pack from the start, or the result is disconnected documents instead of one governable pack.
Taking a polished dashboard as evidence the analytics architecture is sound.
Instead: Settle sources, semantics, authority, metadata, and refresh logic before the visual layer. Inconsistent figures across dashboards or against a regulatory submission point to an ungoverned semantic layer.
Leaving a model that influences real decisions outside the architecture as a data-science tool.
Instead: Once its output changes a decision, the model is an architecture element. Give it an owner, a repository entry, lineage, a registry version, human oversight, and a risk tier that sets how heavy those controls must be.
Core distinctions
- Phase C develops data and application architecture together because authority, integration, and publication choices are coupled, and separating them produces conflict at the boundaries
- An information map groups information into business-relevant domains that survive system change, not an inventory of current databases
- Source of truth is useful only when bounded by entity, attribute, consumer, and purpose, and the five-layer ownership model records the distributed reality
- Customer MDM settles six decisions before any platform: entity definition, shared attributes, authority, matching rules, stewardship, and the decision use case
- The CIM is a shared semantic vocabulary, not a mandatory schema, and alignment means documented, governed mappings at the boundaries
- Metadata carries six minimum concerns, meaning, provenance, stewardship, lifecycle, classification, and quality, and without them technically accurate data can mislead
- A dashboard is an output of the analytics architecture, and the semantic layer beneath it is the most common failure point
- An ABB is technology-aware but product-neutral, and an SBB is product-specific and constrained by the ABB it realises
- The Phase C validation test is end-to-end traceability: a governance board can trace any published figure back to its authoritative source through every layer of the pack
That is the whole stage in one place: the paired domain, the information map, the authority model, customer and utility master data, metadata and the analytics chain, building blocks, integration decisions, the joined-up pack with its traceability test, and models governed as architecture elements. The scenario practice now puts those authority, metadata, integration, and building-block decisions to work on realistic London Grid Distribution situations, one Phase C judgement at a time.
Sources and further reading
- The TOGAF Standard, 10th Edition (C220)Phase C objectives and steps, the ABB and SBB characteristics in Part 4, and the interoperability requirements technique this stage applies.
- G190, Information MappingThe method for structuring enterprise information into domains named in enterprise language.
- G21B, Customer Master Data ManagementThe six decisions on identity, authority, matching, and stewardship that precede any MDM platform choice.
- G234, Metadata ManagementThe six minimum metadata concerns that let consumers interpret and trust published information.
- G238, Business Intelligence and AnalyticsThe four-layer decision-support chain and the analytics patterns behind trustworthy decision support.