Skip to content
Unit 9 of 14
0% complete
RFPLumen DiagnosticsOpen anytime - every module Task uses this brief

Module 7 — Pre-sales, Estimation, Discovery, Construction, Transition

Contents · 35

Engagement models

The engagement model is not commercial trivia — it determines what you are allowed to assume about team stability, scope change and who carries the risk of being wrong. Time pressure during pre-sales often forces a shortcut Ian Gorton calls a marketecture — a one-page, informal depiction of a system's structure and interactions, with just enough labels and boxes to convey the design philosophy behind it. It is not a lesser form of architecture work; it is a discussion vehicle for stakeholders during design, build, review and the sales process itself, and a legitimate starting point for the deeper analysis that follows once the engagement is won. The spectrum, roughly in order of how much the supplier owns:

ModelDescriptionArchitectural consequence
Project-based / defined scopeThe supplier delivers an agreed scope. Reduces cost and time to market by buying expertise and capacity you do not haveWeakens badly when requirements evolve, or when the relationship will outlast one project. Pushes toward up-front design because change is expensive to transact
Dedicated team (often an offshore or nearshore development centre)The client drives requirements; the supplier manages staffing, retention and ramp-up. Accommodates variety — new development, legacy modernisation, maintenance, testingTeam continuity makes incremental architecture viable. You can defer decisions because the people who made them are still there
Managed product developmentThe supplier owns end-to-end delivery: product management, architecture, implementation, operations and supportThe architect owns outcomes rather than documents. Operational qualities — observability, deployability — stop being someone else's problem
Platform partnershipThe supplier operates a platform on a continuing basis rather than delivering a projectDesign for continuous change is the whole job. Total cost of ownership dominates build cost in every decision

Two practical notes. The pricing model is often implied by the engagement model — staff augmentation tends to time-and-materials, defined scope tends to fixed price — and each puts the risk of estimation error on a different party. And architects are rarely involved in contract negotiation, which is usually owned by sales, account and delivery management; but you should know which model you are working under, because it silently sets the design constraints above.

Before an engagement model is even chosen, the client's procurement process usually runs through a named sequence: a Request for Information (RFI) gathers background on candidate vendors, ahead of an RFP or RFQ; a Request for Quotation (RFQ) is used when the specification is already fixed and price is the deciding factor, skipping the back-and-forth an RFP requires; and a Request for Proposal (RFP) — the most formal of the three, with strict rules on content, timeline and vendor response — asks for a proposed solution to a stated problem. Which of these you are responding to changes how much solution detail, WBS and estimate an architect is expected to produce at that stage.

Estimation: art and science

Character
Science of estimationVery mathematically intensive and can be quite accurate. Gives you theoretical numbers. Best implemented by software tools
Art of estimationInvolves heuristics and rules of thumb. Gives you practical help on real projects. Works best when supported by the estimation science

The common idea is that an estimate is a prediction of the project outcome.

DeMarco additionally introduces a quality — or accuracy — attribute of the estimate. According to him the default definition among professionals is "the most optimistic prediction that has a non-zero probability of coming true". He argues a better definition is "a prediction that is equally likely to be above or below the actual result."

Over- vs underestimation

100% accurate estimates are very rare, so if we are going to err, is it better to err on the side of overestimation or underestimation? The penalties are asymmetric:

DirectionPenaltyWhy
UnderestimationNonlinear and unboundedPlanning errors, shortchanging upstream activities, and the creation of more defects cause more damage than overestimation does
OverestimationLinear and boundedParkinson's Law and Student Syndrome. Work will expand to fill available time, but it will not expand any further
  • Parkinson's Law — "work expands so as to fill the time available for its completion". If you give a project team 6 months to complete a project that could be completed in 4, the team will find a way to use up the extra 2 months
  • Student Syndrome — if developers are given too much time they'll procrastinate until late in the project, at which point they'll rush to complete their work, and they probably won't finish on time

In practice, significant overestimation leads to an underestimation-like effect because of artificial increasing of scope. There is an error range which doesn't cause any significant problems: 5%–10%.

Accuracy vs precision

Precision is independent of accuracy. People make assumptions about accuracy based on precision, so the precision you use should match the accuracy of your estimates.

A good analogy: imagine a basketball player shooting baskets. If the player shoots with accuracy, their aim will always take the ball close to or into the basket. If the player shoots with precision, their aim will always take the ball to the same location, which may or may not be close to the basket. A good player will be both accurate and precise — shooting the ball the same way each time and each time making it in the basket.

Inputs, outputs, and the sources of uncertainty

Input information consists of information about the project being estimated and about the capabilities of the organization that will be implementing and delivering it.

  • Project data includes requirements both functional and non-functional, priorities (e.g. quicker is better) and constraints (e.g. the technology stack is already defined)
  • Organizational influences include data about the organization's processes and the personnel's experience and skills

Output consists of estimated scope, effort, schedule and cost.

The estimation procedure itself has to be unbiased and should produce the same output on the same input data. The project data can be adjusted until the estimation process produces an acceptable outcome.

There are two main sources of estimation uncertainty: incomplete or inaccurate input data, and inaccuracies of the estimation process itself.

CauseEffect
Omitted functional and especially non-functional requirementsErrors in project size estimation
Wrong assumptions about personnel's skills and performanceErrors in both effort and schedule estimation
Forgotten activities such as requirement analysisErrors in effort estimation
Insufficient experience in the particular business areaErrors in project size estimation
Unfamiliar technologiesWrong assumptions about performance, and thus errors in effort estimation

The cone of uncertainty

Software development is a process of gradual refinement. You start with a general product concept — the vision of the software you intend to build — and refine that concept based on the product and project goals. Uncertainty in a software estimate results from uncertainty in how the decisions will be resolved; as you make a greater percentage of those decisions, you reduce the estimation uncertainty.

  • The cone represents the best-case accuracy that is possible to have in software estimates at different points in a project
  • The cone doesn't narrow itself. It narrows through project control. If the project is not well controlled, the cone of uncertainty becomes a cloud of uncertainty
  • The schedule variability is much lower than the effort variability, because schedule is calculated as a cube-root function of effort for large projects

Expressing probability

The curve of a probability distribution describes the probability associated with different estimate values; each point represents the chance of the project finishing exactly on that date, or costing exactly that much. What you usually want is the probability of delivering on or before a particular date, or at or below a specific cost or effort.

You can express probabilities in numerous ways:

  • A "percent confident" attached to a single-point number — "we're 90% confident in the 24-week schedule"
  • Best case and worst case, which implies a probability — "we estimate a best case of 18 weeks and a worst case of 24 weeks"
  • A range rather than a single-point number — "we're estimating 18 to 24 weeks"
  • A plus-or-minus qualifier — "6 months, ±1 month"
  • A confidence factor — "there is 80% probability we can deliver by the end of Q4"

The key point is that all estimates include a probability, whether the probability is stated or implied. An explicitly stated probability is one sign of a good estimate.

An essential practice in presenting an estimate is to document the assumptions embodied in it. Then, if the project unfolds in a way that invalidates the assumptions, you can point back to the estimate assumptions as a basis for revising the estimate. The way you communicate an estimate suggests how accurate it is — if your presentation style implies an unfounded accuracy, you lay the groundwork for a difficult discussion about the estimate itself.


Work Breakdown Structure

There are broadly two approaches to WBS breakdown, plus mixed:

TypeDefinition
Deliverable-oriented (product WBS)A hierarchical structure of things that the project will make, or outcomes that it will deliver. It's a classification of the project scope
Task-orientedA hierarchical structure of tasks that, when completed, will result in satisfaction of all project commitments. It's an exhaustive list of work

Questions the course poses for discussion:

  • What are the benefits of the deliverable-oriented WBS? Of the task-oriented WBS?
  • What WBS is potentially more accurate?
  • What WBS leads to easier change of the overall estimate and schedule?

A concrete case to reason about: you need to estimate Reporting, where you will need to build a core reporting framework plus a set of concrete reports. Deliverable. Task. Deliverable.


Estimation techniques

Decomposition is the practice of separating an estimate into multiple pieces, estimating each piece individually, then recombining the individual estimates into an aggregate. Also known as "bottom up" estimation.

Decomposition takes advantage of The Law of Large Numbers: if you create one big estimate, the estimate's error tendency will be completely on the high side or completely on the low side. But if you create several smaller estimates, some errors will be on the high side and some on the low side — the errors will tend to cancel each other out to some degree.

A related, formula-based technique for a single estimate is three-point (PERT) estimation: instead of one number, the estimator states a best case (a), most likely case (m) and worst case (b), which are combined into a weighted estimate — the standard PERT weighting is (a + 4m + b) / 6. It pairs naturally with decomposition, since each decomposed piece can itself be given a three-point estimate before the pieces are recombined.

The four main steps of the estimation process: preparation → estimation of the development efforts → estimation of other activities → finalization.

It is important to estimate non-development activities. Examples: DevOps, business analysis, quality assurance, documentation. You can involve QA, BA and other specialists to help estimate them, or estimate them as a percentage of development efforts.

Conclude by creating the baseline schedule based on your estimated total. You can optionally check your total by using another estimation method.

Always agree what exactly you include in the "Development" activity — the idea is always to be on the same page, because development can include or not include unit testing, integration testing, dev documentation, architectural design, deploy, and so on.

Wideband Delphi

  1. The Delphi coordinator presents each estimator with the specification and an estimation form
  2. Estimators prepare initial estimates individually. Optionally this step can be performed after step 3
  3. The coordinator calls a group meeting in which the estimators discuss estimation issues related to the project at hand. If the group agrees on a single estimate without much discussion, the coordinator assigns someone to play devil's advocate
  4. Estimators give their individual estimates to the coordinator anonymously
  5. The coordinator prepares a summary of the estimates on an iteration form and presents it to the estimators so they can see how their estimates compare with others'
  6. The coordinator has estimators meet to discuss variations in their estimates
  7. Estimators vote anonymously on whether they want to accept the average estimate. If any of the estimators votes "no", they return to step 3
  8. The final estimate is the single-point estimate stemming from the Delphi exercise. Or, the final estimate is the range created through the Delphi discussion, and the single-point Delphi estimate is the expected case

Estimation by analogy

The idea is simple: you create estimates for a new project by comparing the new project to a similar past project.

  • Normally works well in accounts working in a concrete domain and/or with limited technologies and platforms — e-commerce, insurance
  • It is important to use the actual results of the old project and not the old project estimates
  • Can give quick results if you can find a similar project to compare to
  • But the accuracy quickly decreases as the number and magnitude of differences between the projects grows

Story points and #NoEstimates

Approaches to story points: known velocity; defining a transformation coefficient — bad practice; trying to guess velocity; using ideal hours and a load factor.

An ideal man-day, per Henrik Kniberg, is "a perfectly effective, undisturbed day" — Mike Cohn's version of the same assumption is that the story being estimated is the only thing worked on, everything needed is on hand, and there are no interruptions. Because no real day is ideal, you multiply by a load factor to convert ideal man-days into calendar effort; a commonly used default is 0.75, though the right value is account-specific.

The #NoEstimates movement explores alternatives for estimation.

Two rules hidden in the estimation jokes

  1. All estimation techniques are based on previous experience. Model-based techniques use industry-average data; proxy-based (learn-oriented) techniques use the organization's historical data; expertise-based techniques use personal previous experience
  2. Project scope has to be defined as clearly as possible

A practice exercise the course uses: estimating something concrete where students must ask clarifying questions and provide several approaches. The core idea — simply count and calculate; use judgments (e.g. "I think the hall is 70% full") only as a last resort.


Schedule compression

If there is a business need to compress your schedule, keep in mind that shorter schedules require more effort for several reasons:

  • Larger teams require more coordination and management overhead
  • Larger teams introduce more communication paths, which introduce more chances to miscommunicate, which introduce more errors, which then have to be corrected
  • Shorter schedules require more work to be done in parallel. The more work that overlaps, the higher the chance both that one piece of work will be based on another incomplete or defective piece, and that later changes will increase the amount of rework

There is also a limit to how much you can compress your baseline schedule.


Discovery

A high-level picture of possible participants exists, but the list can vary depending on actual needs. There are major roles plus a few more, and normally a coordinator is an Account Manager. The list of responsibilities is not exhaustive and depends on the account or project.

  • The discovery approach should be defined and communicated to the client. A sample: 6 weeks, two streams involved, on-site/offsite presence varies
  • In order to get valuable deliverables, the team should define and communicate to the client the assumptions (e.g. schedule, availability of key stakeholders) and the technical prerequisites
  • A sample-template of an on-site agenda is provided
  • Important: define and confirm with the client a set of expected outcomes, for example:
    • A report including key issues and recommendations, if this is brownfield
    • A roadmap
    • Solution design and estimate

QAW, or any of its lightweight versions, can be used at this stage — details in Module 3.1.


Construction

The idea of the Construction module is to cover the different areas solution architects can be responsible for or involved in during a project implementation phase. Clients and projects are different, therefore SA activities at this stage vary.

To the common question of what an architect should do during implementation, the generic but real answer: we work for and together with our customers, so we as architects must do everything from our side to ensure that projects are completed on time and with the required quality.

In a typical agile project the architect is involved in:

  • Initial design during the inception phase — the Initial Architecture Vision
  • Estimations
  • Up-front design for the next iterations/releases
  • System redesign
  • Design of new or complex functionality

Note: at SA L1 level, treat architecture governance mostly as implementation governance — following the best practices, doing architecture and code review, writing useful documentation. Architecture governance as part of an EA framework is a somewhat different thing and out of scope.

A useful model here is Lencioni's three virtues of an ideal team player, which map unusually well onto architecture work:

QualityA team player who is…
HungrySelf-motivated and hard-working; always thinks about the next step and next opportunity to contribute; always looking for more to do, more to learn and more responsibility; never wants to be considered a person who avoids work or effort
HumbleLacks excessive ego and is not concerned with their status; quick to point out contributions of others and slow to seek attention for themselves; puts the team over self; defines success collectively rather than individually
People SmartHas common sense about people; actively listens to others and stays engaged in conversations; asks good questions; has good judgement and intuition about the subtleties of group dynamics; understands the impact of their words and actions

There are different approaches to the test pyramid; the one presented is just one of them.

Rendering diagram…

Each layer up trades speed and count for realism — many fast unit tests at the base, a handful of slow manual passes at the top.

Transition itself is not one activity but three distinct handoffs, each with its own checklist: transition to production (UAT) — a hard code freeze admitting only critical fixes, user-acceptance and performance testing, and training material for end users; transition to another team or vendor — technical knowledge transfer plus architecture, documentation, process and UX review; and transition to support/maintenance — knowledge transfer and a documentation review and update. Treating transition as one generic wrap-up phase is how architects under-plan it, since each variant has a different audience and a different failure mode if skipped.


Hardware selection

Primary resources are the typical resources that matter for most types of solutions. Additional resources are important for special scenarios:

  • GPU grids are often used for video encoding and for deep learning (neural networks). There are cases where neural network training performance on 1 GPU node (AWS P2) was the same as on a cluster with 144 CPUs — by replacing CPU nodes with GPU nodes the solution became 20× more cost effective
  • Disk size is important for large DBs, NoSQL and big data solutions
  • Physical size/dimensions and energy consumption are critical in IoT and wearable devices

Normally, software performance is bound to one or more primary hardware resources. Knowing which types of workloads and solutions are CPU-bound or memory-bound helps make initial estimations or pick starting points for performance testing.

CPU

ConceptDetail
CPU boundImproving the CPU — faster, more cores, better cache — will improve application performance; the application spends the majority of its time using the CPU doing calculations
I/O waitsThe application is not using CPU and is waiting for I/O — reading from disk, making a request over the network to a DB or service
Cores vs threadsIf your application doesn't have any I/O waits, which is rare, optimal performance is when the number of threads equals the number of cores
Context switchingHaving more running threads than cores leads to interrupting one task and switching to another, which adds CPU overhead. With no I/O wait, context switching degrades performance. When there is I/O wait it is totally acceptable and common practice to have more threads than cores
MetricsCPU usage and load average are the typical metrics. CPU utilization should normally be below 70–80%, and load average — tasks for CPU waiting in queue — below 0.7 per core

Memory

Memory on the node is consumed not only by your application — there are other consumers such as the kernel, page cache, daemons and agents. It's not uncommon to leave extra free memory unallocated to leave more space for page cache, especially in disk-I/O-heavy applications like databases and NoSQL data stores.

  • Adding more memory is not the right way to solve memory leaks — address it in your application instead
  • Swap can extend memory with disk space but slows performance when used; therefore in cases where high performance is required, swap should be turned off
  • The Linux kernel OOM Killer will kill your process if it considers it too aggressive in memory consumption. Ensure proper hardware and app allocated memory sizing, and keep memory utilization below 70–80%
  • Garbage collection performance — allocating too much memory to a process using a GC-backed runtime (JVM, CLR, JavaScript, Go) can and normally will impact performance. Ideally keep nodes smaller and spread memory across the cluster. Depending on the runtime you may want to keep RAM allocated below 8–16 GB, otherwise investment in GC tuning may be required

Disk

MetricDefinitionMost critical for
IOPSInput/output operations per second — the number of IO operations a disk can perform per second, disregarding the size of each operationOperational storages based on RDBMS or NoSQL
ThroughputIOPS × IO size — the volume of IO operations handled per unit of time (GB/s)Big data and analytical solutions
LatencyHow fast a single IO operation can be performed (ns, ms)Low latency systems

Different storage is optimized for different types of performance, so the choice depends on the quality attributes of the system.

  • SSD vs HDD — SSD often provides better performance, especially latency on random reads/writes. The difference on sequential access patterns is less significant. The primary drawback of SSD is cost
  • Local attached disks vs network storage (SAN, NAS) — network storage provides a lot of benefits especially from a maintenance and reliability perspective, and in most cases better TCO. At the same time, using network storage removes "data locality", which is critical for data-intensive applications, increasing disk operation latencies and network I/O
  • Access patterns — random vs sequential, read-intensive vs write-intensive vs mixed. Disks have different effectiveness with different access patterns (see the HDD vs SSD example), so understand your solution's access patterns to make a proper disk selection
  • Security — data security at rest is critical for most solutions. While software encryption is often sufficient, some solutions require advanced security either forced by regulations like FIPS or by protecting sensitive information. In those cases, disks with hardware-based security may need to be considered — with support of SED and ISE

Network

Network is a huge topic deserving its own deck; these are the key points covering the needs of most solutions.

MetricDefinitionCritical for
BandwidthThe max rate of data transferred (MB/s)—
ThroughputThe actual data transfer rateLarge-volume data transfer solutions: big data, video streaming, some ML scenarios like neural networks on distributed GPU grids
LatencyThe delay between sender and receiver first byteLow latency solutions

Reliability — jitter, error rate, network partition — has significant impact on certain types of solutions:

  • Networks with high jitter are not suitable for low-latency solutions: despite "advertised" network performance that might provide sufficient latency in theory, high jitter will cause it to regularly breach SLAs
  • Network partition rate is critical for data stores that are on the CP side of CAP but at the same time try to maintain high availability. Google Cloud Spanner is a CP data storage that maintains 99.999% due to the high quality of Google's network, where network partitions are extremely rare

Software-defined networking (SDN) built on top of network hardware topology can have huge impact on performance and reliability. For example, Google's Andromeda SDN upgrade from version 2.0 to 2.1 reduced network latency for intra-zone VM communication by 40% while using the same hardware.

Security — network security is critical for most modern solutions, and this is one of the biggest challenges in the IoT field. While for most solutions generic cloud network hardware security plus correct network topology (DMZ, VPCs) plus software security (encryption) is enough, additional hardware options exist such as Hardware Security Modules (HSM).

Sizing for availability and geography

Multi-region requirements are driven not only by availability and DR SLA but by the geographical distribution of consumers. At the same time the number of availability zones is purely availability related.

Going below the minimal number of nodes will breach the availability SLA — therefore if that number of nodes provides more capacity than your application needs, consider downsizing the nodes rather than reducing their count.

Autoscaling delay

There is a delay. You normally don't want to increase the number of instances in a cluster the moment utilization reaches a threshold (e.g. CPU > 80%), as it could be an occasional spike. The common approach is to set autoscaling to kick in after the threshold was passed for a few sequential minutes. Bear in mind that once new nodes are added, there is extra start-up time for your app, ranging from milliseconds to minutes depending on its type.

$$\text{autoscaling delay} = \text{autoscaling time threshold} + \text{app startup time}$$

In some cases stress arresting is required to ensure your app can survive big spikes until autoscaling kicks in. With that said, delay in autoscaling is not always suitable for low-latency applications, as typical stress-arresting techniques break latency SLAs — those scenarios require preemptive scaling.

For distributed datastores that leverage sharding or partitioning, autoscaling is rarely used, because changing the topology of a partitioned cluster often requires re-partitioning before a node can start serving requests. On large volumes of data this may take hours, and the load spike can be over by then. At the same time, for non-partitioned data stores, creating additional read replicas is a much faster process and can be considered for autoscaling.

Sizing tactics

Estimation tactics are usually employed in the initial stages of the project, when it is not possible to test whether a hardware configuration will correspond to the solution's quality attributes. Often used as a starting point for the first performance testing.

TacticExample
Similar systems"We had a previous version of our service running on 15 c4.large nodes; as we didn't make any significant changes that might affect performance, we will start with the same cluster"
Previous experienceExamples of CPU/RAM/disk/network-bound systems that give you an idea of what type of resource would be most needed for your solution
Calculation basedKnowing the number of rows in your DB and their size, you can estimate the total data volume. Knowing the size of request/response and max connections to your service, you can estimate the RAM required
Vendor recommendationsUnless it's a fully custom-built app, most software products (DBs, NoSQL) have recommended hardware requirements. Remember it is not possible to provide recommendations for every possible use case — never trust them as-is
Zooming / proportional testingIf developing a bank system, measure performance on 1K transactions per 100 users, then propagate the results to the ASRs accordingly

Fact-based tactics are based on monitoring of a real implementation as close as possible to production, often peak, loads.

TacticExample
Current systemWhether it is a modernization effort or a cloud migration, whenever there is an existing system, collect as many metrics as possible. Not only may some hardware requirements stay the same, it will also help with estimations
Performance testingCritical for hardware sizing — make sure you test for peak loads as well
Shadow trafficIf replacing an existing implementation, before releasing to production it is not uncommon to route production traffic to it so traffic goes to both old and new solutions, while the old one still serves production responses. Not a replacement for performance testing, as it might not show behaviour under peak load — but the benefit is that it is real traffic, as opposed to the often synthetic tests used in performance testing
Production monitoringOnce the system is in production, monitor it closely to see if you under- or over-provisioned hardware resources, and adjust accordingly
Cloud cloneUse cloud infrastructure to test a full clone of the system — rent cloud capacity for a short period on a separate account, monitor performance and VM loads, then estimate the resources actually used
"The best one minus one"A top-level CIO tactic: when customers go for hardware they are OK to pay for good enough, but not top-notch systems. So select the line-up above the average — not Intel Core i9, but i7

Which tactics you apply depends on what information you have or can get. If you have an existing system and are replacing it, you can use the same RPS adjusted with usage growth projections and other factors. If this is a completely new system you'll need to get creative and estimate the number of concurrent users and their behaviour. If it is a service consumed by other existing services, you can look at their RPS.

Most distributed data platforms publish minimum and recommended production cluster specifications — node class, memory, disk type and node count. Use them as the starting point for performance testing and adjust from measurement, never as the answer.

Worked performance-testing iteration

A simple microservice that just serves GET requests. Expected peak load 3,000 requests per second; SLA: response latency percentile 95 shouldn't exceed 20 milliseconds.

StepApproachOutcome
1. Start smallStart with a small hardware sizing you believe makes sense based on experience or an estimation-based technique. Start with small loads too and increase gradually to see where the system breaks or slows downIn both cases the target 3k RPS was not reached, CPU utilization was high, and the latency SLA was breached
2. InterpolateFrom the observations CPU was obviously the bottleneck, so increase both the number of CPUs per node and the number of nodes. Also, while memory wasn't bottlenecked it constantly held at 80% utilization, close to max, so add memory as wellThis setup reached the required 3k RPS without issues. But CPU utilization at peak was 22% — it looks like we over-provisioned
3. OptimizeReduce the size of the cluster, since we over-provisionedUnder-provisioned this time, but not as badly as on the first run. Pretty close
4. Final resultIncrease cluster size one more timeFound the happy medium: CPU utilization at 53% under peak load, which is safe even if the workload is slightly higher than expected peaks

Usually performance testing takes much more iterative steps than in this example, especially where your service is backed by a data storage where you have impact on hardware sizing as well — in that case you have much more variability of parameters to test.


APM and TCO

APM (Application Performance Monitoring) has a typical flow of analysis through the solution.

TCO — as solution architects we must understand what TCO means and why the business cares. It is life-cycle cost analysis, not merely build cost, and the dominant term is maintenance. Calculating a TCO figure rather than gesturing at one means defining the scope and length of ownership, building a cost model as a matrix of resources against cost type, and summing three categories per resource: acquisition cost (license, hardware, training, initial development), operating cost (keeping the resource running — updates, support, training) and change cost (enhancements and bug fixes). That structure is what lets an architect justify a design decision in cost terms instead of opinion.

How much of a system's lifetime cost is maintenance? Software engineering literature has put it at roughly 60–80% of total lifetime cost for decades, and studies vary mainly in how they draw the boundary. The corollary for an architect is uncomfortable: most of the money is spent after you have gone, and a large share of it is spent working around decisions made early.

The classic decomposition comes from Lientz and Swanson — corrective, adaptive and perfective maintenance — with preventive added later by ISO/IEC 14764. Typical shares, which vary by study but not in their ordering:

Maintenance typeWhat it isRough share
CorrectiveFixing defects found after deployment~20%
AdaptiveChanging the system so it stays effective as its environment changes~25%
PerfectiveImproving performance or maintainability without changing behaviour~5%
EnhancementNew capability — continuing functional evolution~50% or more

Read that table next to the Module 3 modifiability material and the conclusion is uncomfortable but useful: the majority of what you will spend on a system after release is not fixing it — it is changing it. Which is precisely why binding time, cohesion and coupling decisions dominate TCO, and why "we'll make it flexible later" is the expensive option.


Worked estimate — Lumen Diagnostics

The Module 7 exercise applied to the case study. The value is in the stated rules and the traceability, not the numbers.

Ground rules that separate an estimate from a guess

  • State what is included — core development, unit tests, integration tests, documentation — and what is not — business analysis, support, training
  • Separate by category — front end, back end, data, operations. This separation is what makes a resource plan possible at all
  • Decide deliberately whether architecture effort sits inside development or is estimated as its own line
  • Discovery and transition belong in the plan. If your estimate excludes them, say so explicitly rather than letting the omission pass silently
  • Non-development work — management, business analysis, QA — is commonly estimated as a percentage of development effort, which is defensible only if you state the percentage and why
  • Every number must reconcile with the resource plan. If development totals 100 person-days, roughly 100 days of development roles must appear in the plan. A total that does not reconcile is a total nobody checked
  • For the discovery phase, produce an activity grid — days across, sessions down, each cell naming the activity and who must attend

The percentage rules, and why they are a judgement

RuleTypical rangeWhat moves it
QA effort as a share of development20–60%A system with a hard performance SLA needs a performance specialist as well as functional testing, which is what pushes it toward the upper end. A straightforward CRUD portal sits near the lower end
Management as a share of total team effort~10%Team count and number of parties to coordinate
Operations engineeringSpans the whole deliveryNot a phase. Environments, pipelines and observability are needed from the first sprint to the last

Pick a figure, write down the reasoning, and defend it in review. A ratio quoted without reasoning is the most common way an estimate becomes indefensible under pressure.

Shape of the work breakdown

Deliverable-oriented, with days per role (SA · BA · Front end · Back end · Data · Ops · QA):

PhaseDeliverableRepresentative activities
DiscoveryOnboardingReview the referral-management and identity integrations; assess partner-practice capability and data quality
WorkshopsStakeholder interviews, a quality attribute workshop, requirement and constraint refinement and sign-off
FinalisationProduce the architecture document: application, data and integration architecture
ImplementationFoundationsProvision environments; establish the delivery pipeline; define and document access roles; set up the documentation space
Architecture and designApplication architecture, data architecture, integration architecture, per-region deployment topology
Platform capabilitiesWorkflow configuration service; offer-and-confirmation service; availability search; notification service; logging, audit and monitoring
PortalsPatient portal; administration portal; partner portal; each with its own API surface
IntegrationsReferral-system synchronisation; identity service; single sign-on; the one automated partner API; manual upload path
TestingFunctional, integration, performance and residency-compliance testing
TransitionReadinessProduction configuration, data migration and reconciliation, runbooks, support handover
LaunchPhased regional rollout with rollback plan

Roll the days up per role, apply a rate per role, and the resource plan produces the cost. That final step is only possible because the estimate was separated by category in the first place — which is why that ground rule is not bookkeeping.

Worked cloud cost model

The Module 9 exercise asks you to calculate running cost. The method matters more than the arithmetic: derive call volumes from user behaviour, then apply the price list. Never the reverse — a spreadsheet that starts from prices tells you nothing about what to change.

Step 1 — the load drivers

DriverValueAssumption
Named users (staff, partners, patients with accounts)52,0008% concurrent → ~4,160 concurrent sessions
Appointment and study records6.5 million12% updated monthly, 4% added → 1.04M record changes per month
Search index size~40 GBDrives node sizing, not call volume

Step 2 — derive per-service call volumes, showing the arithmetic

ServiceMonthly callsDerivation
Scheduling orchestrator36,400,0001.04M changes × 5 calls per stage × 7 workflow stages
Workflow configurator7,280,0001 call per stage per changed record → 1.04M × 7
Partner availability sync18,200,000Estimated at half the orchestrator volume
Availability search15,200,00025% of concurrent users search 4×/hour → 4,160 searches/hour × 5 calls × 730 hours
Notification service4,160,0001.04M changes × 4 calls per notification
Reporting2,000,0008% of users run reports monthly, 40 reports each, 12 calls per report
Identity and user management250,000Low volume, dominated by session establishment
Patient feedback150,000Small share of appointments, plus browsing
Monitoring5,260,00012 services × 60 checks/hour × 10 calls × 730 hours
Logging and audit5,260,000Comparable to monitoring
Data access layer41,820,00050% of business-service call volume
Total≈ 136,000,000

Step 3 — apply the price list

Illustrative published list prices; substitute your provider's current rates.

ItemBasisUnitsMonthly
API gateway$3.50 per 1M calls136.0M$476.00
Functions — invocations$0.20 per 1M136.0M$27.20
Functions — duration$0.00001667 per GB-second17.0M GB-s$283.39
Notifications — email (80%)$2.00 per 100,0003.33M$66.56
Notifications — SMS (15%)$6.00 per 1,000624,000$3,744.00
Notifications — push (5%)$0.60 per 1M208,000$0.12
Search cluster$0.22 per node-hour2 nodes × 730 h$321.20
Search storage$0.11 per GB-month40 GB$4.40
Total≈ $4,923 / month (~$59,000 / year)

Three things to take from this, all of them architectural rather than financial:

The SMS channel is 15% of notifications and 76% of the entire bill. Unit economics beat volume intuition. That single line justifies a design conversation — make SMS opt-in, or reserve it for the stages where a missed message costs an appointment — and no amount of tuning elsewhere recovers comparable money.

The search cluster was sized from the index, not from traffic. 40 GB of index chose the node class; the call volume merely confirmed two nodes suffice. That is the "calculation based" sizing tactic applied literally.

The data access layer is 31% of all calls and is pure overhead. It appears in no requirement. It exists because the design put a layer there — which makes it exactly the kind of hop the availability arithmetic in Module 6 says to question.

Exercises — Module 7

Task. Plan delivery for Lumen Diagnostics. Output: WBS, effort estimate, resource plan, discovery plan (at least one week in detail). Numbers must reconcile.

  • Create a WBS (delivery-based, task-based, or mixed)
    • Mixed WBSs are easy to get wrong; keep the two kinds distinct
  • Estimate effort, split by category
    • Split: front end / back end / data / operations
    • In scope: core development, unit tests, integration tests, documentation
    • Out of scope: business analysis, support, training (unless you explicitly include them)
    • Say whether architecture effort sits inside development
    • Discovery and Transition: in the plan, or explicitly excluded
    • Non-development (management, BA, QA) as a stated percentage of development is fine
  • Resource plan for all main roles
    • Must reconcile with the estimate
  • Assumptions and risks
  • Discovery plan (a day-by-role grid, not a slogan)
    • Timeline
    • Activities and roles (who attends QAW, playbacks, integration workshops)
    • Who does what, day by role
    • Expected outcomes
    • At least one week in detail