Style vs pattern vs tactic
There is no single definition and no clear separation between styles and patterns — as usual, everything depends on the usage context.
| Term | Meaning |
|---|---|
| Architectural style | The highest level of granularity in architecture. Specifies layers, high-level modules of the application, how those modules and layers interact, and the relations between them. A general direction for how you plan to build the solution — "we are going to use event-driven design" |
| Architectural pattern | Solves problems related to the architectural style, becoming more specific because it takes the context into account. For example: "what classes will we have and how will they interact in order to implement a system with a specific set of layers", or "what high-level modules will we have in our service-oriented architecture and how will they communicate", or "how many tiers will our client-server architecture have" |
| Architectural tactic | Design decisions that improve individual quality attribute concerns. Tactics implemented in existing architectures can significantly impact the architecture patterns in the system; tactics selected during initial design significantly impact which patterns are used and how they must change to accommodate the tactics |
A design solution for a concrete context and problem is an architectural pattern. A pattern includes or consists of a number of architectural tactics.
Foundational principles
| Principle | Definition |
|---|---|
| Single Responsibility (SRP) | Every object should have a single responsibility, and all its services should be narrowly aligned with that responsibility. At some level cohesion is considered a synonym for SRP |
| Separation of Concerns (SoC) | The process of breaking a program into distinct features that overlap in functionality as little as possible. A concern is any piece of interest or focus in a program; typically concerns are synonymous with features or behaviours |
| Bounded Context | A central pattern in Domain-Driven Design and the focus of DDD's strategic design, which is all about dealing with large models and teams. DDD deals with large models by dividing them into different bounded contexts and being explicit about their interrelationships. A Context Map is the global view of the application as a whole; each bounded context fits within it to show how they should communicate and how data should be shared |
| Principle of Least Knowledge | Any component or object should not have knowledge about the internal details of other components. This avoids interdependency and helps maintainability |
Coupling
A software unit may be of any granularity: methods in a class, classes, subsystems, packages, modules, services, systems. Coupling is the degree of interdependence between software modules. Two tightly coupled modules are strongly dependent on each other; loosely coupled modules are not dependent; uncoupled modules have no interdependence at all. A class with high (strong) coupling relies on many other classes.
Types of coupling, strongest to weakest:
- Classes mutually access each other's private data — a very strong form; you can no longer change one class without considering the other
- Classes communicate via a global data structure — direct dependencies are released and outsourced to the global structure, but coupling is still very strong: all changes affecting the global data also affect all classes working with it
- Classes communicate only via method parameters — considerably lower coupling; the methods involved contain only essential data, so changes cause only local changes to the relevant methods
- No coupling — a system of connected objects communicating via messages
Example: for loosely coupled classes, changing something major in one class should not affect the other. High coupling makes code difficult to change and maintain: because classes are closely knit, a change could require reworking an entire system.
Loose coupling can be achieved by providing an interface other classes use rather than calling a specific implementation directly. When the app calls IDatabase.exportToFile(), you can change the underlying DB from Oracle to MySQL without changing the calling class's code.
A real drawback story from the homework: an application was tightly coupled with a third-party service. Performance testing — and functional testing — was painful, because when the third-party tool was down or experiencing temporary degradation it affected test results. The fix was to talk to an interface and build a stub implementing it.
Cohesion
Cohesion refers to the degree to which the elements inside a module belong together — it measures the strength of relationship between pieces of functionality within a given module. In highly cohesive systems functionality is strongly related. Cohesion is an ordinal measurement, usually described as "high" or "low".
Modules with high cohesion are preferable, because high cohesion is associated with robustness, reliability, reusability and understandability. Low cohesion is associated with being difficult to maintain, test, reuse or even understand.
Low cohesion example: methods in a class doing something totally unrelated to each other — a UTIL class serving as a Swiss army knife providing file export, string operation helpers, FTP connectivity and whatever else. The solution is to separate the helper classes.
Cohesion is the indication of the relationship within a module. Coupling is the indication of the relationships between modules.
A related design axis is whether a component is stateful or stateless. A stateful component keeps state shared across several requests, which transfers less data per call but forces you to deal with synchronising concurrent requests against that shared state; a stateless component puts all the information a request needs into the request itself, trading a larger per-request payload for the ability to route any request to any instance without coordination — which is why stateless components are the easier default to scale horizontally.
Monolith
A monolithic architecture suits simple, lightweight applications for POC or MVP purposes. But one major drawback is tight coupling — over time monolithic components become tightly coupled and entangled, which affects management, scalability and continuous deployment. Other cons stemming from tight coupling:
| Con | Detail |
|---|---|
| Reliability | An error in any of the modules can bring the entire application down |
| Updates | Due to a single large codebase and tight coupling, the entire application must be deployed for each update |
| Technology stack | A monolithic application must use the same technology stack throughout; changes are expensive in both time and cost |
A real drawback story: a desktop application compiled and deployed as one huge executable, used by various departments. Each deployment affected all departments, and there had to be at least core regression testing for every department's functionality after a change in even one particular place.
Layered
One of the powerful features of the layered pattern is the separation of concerns among components.
Rules and their nuances:
- Normally requests always go from up to down. A lower layer should never perform a request to upper layers
- But a layer is allowed to make upward calls as long as it isn't expecting an answer from them — this is how the common error-handling scheme of callbacks works
- Requests from A to C are allowed to bypass B. In that case layer B is considered an open layer, and you take away the benefits of having isolated layers. The Open/Closed principle allows skipping some layers intentionally
- Any set of boxes stacked on top of each other does not constitute a layered architecture. If everyone is allowed to use everything, it is not layered
- The key provides the answer to "what allows to use what"
Onion architecture consists of typical layers, but it is not obvious.
For more depth: search for "layer bridging" in Software Architecture in Practice by SEI, and read chapter 1 of Software Architecture Patterns by Mark Richards.
The figure the Module 4 Task refers to. Layers A (top), B (middle), C (bottom). Solid arrows are allowed in a closed layering; the dashed arrow is an open layer (A skips B); the red arrow is an upward call that expects an answer, which is not layered.
A pragmatic view: stick to the classic approach — A talks to B only, B talks to C only — as much as possible. Jumping from A to C might be needed when performance isn't good enough and you need to speed up; that can be a valid trade-off. Talking backwards from C to B is hard to justify and looks like something that could not be called layered architecture at all.
A related but distinct style is module-based architecture, which treats modularity as a first-class deployment unit rather than a set of call-direction rules: each module declares a name, a version and the dependencies it requires, and hides its implementation behind an explicit contract, so modules can be versioned, replaced and loaded independently at design time and runtime. Because plain objects, packages and JAR archives don't enforce this kind of boundary, platforms such as OSGi and the Java Module System (Jigsaw, since Java 9) exist specifically to let the runtime check and enforce those module contracts.
SOA and microservices
Scope is the difference: Service-Oriented Architecture is enterprise scope; microservices architecture is application scope. In SOA, reusability of integrations is the primary goal, and at an enterprise level striving for some level of reuse is essential — reusability and component sharing increase scalability and efficiency.
Microservices are a SOA instantiation driven by DevOps practices, emphasising CI/CD: smaller services, smaller responsibilities, less coupling — more infrastructure mess.
Nine characteristics:
- Component-based architecture — components are good again because of low coupling, independent deployment and scalability
- The monolith is split according to business functions, not according to organisational structure
- Treat software not as work to do and hand out to maintenance, but as a product to be developed over its lifetime — this emphasises business value for users
- Simple communication with no logic — no complex routing or transformation as in an ESB — plus a smart service that contains the logic inside
- In microservices you can take advantage of the different tools and approaches that better fit a task. Responsibility is distributed as well, and this impacts quality
- Instead of an org-structure- and vendor-licensing-driven approach to storing data, microservices use separate DBs per service — polyglot persistence — which is more efficient because you use the proper technology for a task
- Extensive use of automated infrastructure platforms (e.g. AWS) for both deployment and operations
- Design software tolerant to failures as much as possible — monitor, restore
- Promotes making changes and evolving the system, since you can touch a component with no impact on other components
A reference microservices layering:
| Layer | Contents |
|---|---|
| Consumer | Shows the different clients of the product |
| Delivery | Balances requests, caches static web content, addresses requests to the right application |
| Aggregation | Server-side UI application, and different sets of APIs for specific clients which aggregate business services |
| Service | Business services, system services, gateway and discovery services |
| API Gateway & Service Discovery | Internal components providing routing, discovery, balancing and security for the service layer |
| Business Microservices | Implement business logic and processes |
| System Services | Existing services in the customer's environment — LDAP, printing service, email server, time service |
| Platform Services | Provide support for security and microservices management |
| Infrastructure | Provides virtual infrastructure for the platform |
A microservice is an architecture that structures the application as a set of loosely coupled, collaborating services.
Challenges and drawbacks. It can be very hard for a small company to maintain the whole lifecycle, because it requires extra methods and tools to support the development process — you need DevOps tools such as CI/CD servers, configuration management platforms and APM tools to manage the network. An additional drawback is performance: sending messages back and forth between microservices comes with a certain overhead. The most challenging part is finding how to partition the solution. Orchestrating numerous development teams is also more difficult.
Event-driven architecture
The event-driven pattern is a popular distributed asynchronous architecture pattern used to produce highly scalable applications. It is also highly adaptable and can be used for small applications as well as large complex ones. It is made up of highly decoupled, single-purpose event-processing components that asynchronously receive and process events.
Use it when you want to achieve a highly decoupled, asynchronous and distributed architecture. Once you achieve a high degree of decoupling you can scale architecture components independently, which makes it a good option for modern, distributed, cloud-enabled applications that are horizontally scalable and resilient to failure.
Most often just a part of a system — some sub-system or component — is implemented according to the event-driven style.
Mediator topology — commonly used when you need to orchestrate multiple steps within an event through a central mediator. It is useful for events that have multiple steps and require some level of orchestration to process.
Four main component types: event queues, an event mediator, event channels, and event processors. The flow starts with a client sending an event to an event queue, which transports it to the mediator. The mediator receives the initial event and orchestrates it by sending additional asynchronous events to event channels to execute each step of the process. Event processors listen on the channels, receive the event from the mediator, and execute specific business logic.
- The event-mediator component is responsible for orchestrating the steps contained within the initial event
- Event channels are used by the mediator to asynchronously pass processing events related to each step to the processors; channels can be message queues or message topics
- Event processor components contain the application business logic necessary to process the processing event
The event mediator can be implemented in a variety of ways, and as an architect you should understand each option to ensure the solution matches your needs.
Broker topology — differs in that there is no central event mediator; instead the message flow is distributed across the event processor components in a chain-like fashion through a lightweight message broker (ActiveMQ, HornetQ). Useful when you have a relatively simple event processing flow and do not want or need central event orchestration. Each event-processor component is responsible for processing an event and publishing a new event indicating the action it just performed. Channels within the broker can be message queues, message topics, or a combination.
Related: the Actor model is also covered as an event-based approach. In this model, everything is an actor — a computational entity that, on receiving a message, can concurrently send a finite number of messages to other actors, create new actors, and designate the behaviour it will use for the next message it receives, with no assumed ordering between these actions. Decoupling the sender from the recipient's location and lifecycle is the model's fundamental contribution, and it is what lets actor frameworks (Akka, Erlang/OTP, Orleans) implement asynchronous, highly concurrent processing without shared mutable state or locks.
REST and the Richardson Maturity Model
The model actually starts two levels below hypermedia: Level 0 (the "Swamp of POX") uses HTTP purely as a transport for RPC-style calls against a single endpoint, with no real use of the protocol; Level 1 introduces individual resources, each with its own URI, in place of that one endpoint; and Level 2 adds proper use of HTTP verbs and status codes (GET, POST, PUT, DELETE) to operate on those resources. Only once all three are in place does level 3 add hypermedia controls on top.
Level 3 of the model is hypermedia controls. It is worth reaching when the client doesn't know the full REST API specification, or when the client's behaviour depends on internal logic implemented inside the REST service.
Example: when a client GETs an account balance it also receives a list of possible actions as links. Per the service's internal logic, when the balance is negative the only available action is "deposit money"; otherwise deposit, withdraw and close are offered. The client may also change UI elements accordingly, disabling or hiding buttons.
A big advantage of RESTful APIs that have attained level 3 is the ability to layer and apply cache constraints. Hypermedia helps customise content for new environments while retaining UX across the board: by including unique URLs within a response package, hypermedia APIs tell clients what capabilities are possible and in what scenarios. Hypermedia links can immediately reflect new user permissions without breaking changes on the client side. Hypermedia is a way for APIs to respond to next-generation platforms and the new issues that arise — beneficial in messaging applications such as Slack.
CQRS
Traditional CRUD disadvantages:
- It often means a mismatch between the read and write representations of the data, such as additional columns or properties that must be updated correctly even though they aren't required as part of an operation
- It risks data contention when records are locked in a collaborative domain where multiple actors operate in parallel on the same data, or update conflicts caused by concurrent updates under optimistic locking. These risks increase as complexity and throughput grow. The traditional approach can also negatively affect performance due to load on the data store and data access layer, and the complexity of queries required
- It can make managing security and permissions more complex, because each entity is subject to both read and write operations, which might expose data in the wrong context
When CQRS suits. Very useful in the case of large differences between the numbers of read and write operations — social networks, for instance. You can scale both sides independently to achieve better I/O performance and support parallel operations on the same datasets. Even without a big disparity, you can apply different optimization strategies to the two sides — for example using different database access techniques for read and update. It is particularly relevant for handling high-performance applications, and for writing normalized data very quickly then denormalizing it so reads read prepared data with no need to join multiple tables.
Considerations before implementing:
- Dividing the data store into separate physical stores for read and write can increase performance and security, but adds complexity in resiliency and eventual consistency. The read store must be updated to reflect changes to the write store, and it can be difficult to detect when a user has issued a request based on stale read data — meaning the operation can't be completed
- Apply CQRS to limited sections of your system where it will be most valuable
- A typical approach to deploying eventual consistency is to use event sourcing in conjunction with CQRS, so the write model is an append-only stream of events driven by command execution, and those events update materialized views acting as the read model
Not recommended for implementation across the whole system. There are specific components of an overall data management scenario where CQRS is useful, but it adds considerable and unnecessary complexity when not required. Be aware it can be complex to implement and provides eventual consistency only.
Event sourcing
Events are immutable and can be stored using an append-only operation. The user interface, workflow or process that initiated an event can continue, and tasks handling the events can run in the background. Combined with the fact that there is no contention during the processing of transactions, this can vastly improve performance and scalability, especially for the presentation level.
Events are simple objects describing an action that occurred together with any associated data required. They don't directly update a data store — they're simply recorded for handling at the appropriate time, which simplifies implementation and management.
Events typically have meaning for a domain expert, whereas object-relational impedance mismatch can make complex database tables hard to understand: tables are artificial constructs representing the current state of the system, not the events that occurred.
Event sourcing can help prevent concurrent updates from causing conflicts because it avoids directly updating objects in the data store. However the domain model must still be designed to protect itself from requests that might result in an inconsistent state.
The append-only storage provides an audit trail usable to monitor actions taken against a data store, regenerate the current state as materialized views or projections by replaying the events at any time, and assist in testing and debugging. The requirement to use compensating events to cancel changes provides a history of changes that were reversed — which wouldn't be the case if the model simply stored current state. The list of events can also analyse application performance, detect user behaviour trends, or obtain other useful business information.
The event store raises events and tasks perform operations in response. This decoupling of the tasks from the events provides flexibility and extensibility: tasks know about the type of event and the event data, but not about the operation that triggered the event, and multiple tasks can handle each event. This enables easy integration with other services and systems that only listen for new events raised by the event store. However, event sourcing events tend to be very low level, and it might be necessary to generate specific integration events instead.
When to use it:
- When you want to capture intent, purpose or reason in the data — changes to a customer entity captured as specific event types such as Moved home, Closed account, or Deceased
- When it's vital to minimize or completely avoid conflicting updates to data
- When you want to record events and be able to replay them to restore state, roll back changes, or keep a history and audit log — for example when a task involves multiple steps and you need to revert updates then replay some steps to bring data back to a consistent state
- When using events is a natural feature of the application's operation and requires little additional development effort
- When you need to decouple the process of inputting or updating data from the tasks required to apply those actions — to improve UI performance, or to distribute events to other listeners. For example, integrating a payroll system with an expense submission website, so events raised by the event store in response to website updates are consumed by both the website and the payroll system
- When you want flexibility to change the format of materialized models and entity data if requirements change; or — with CQRS — when you need to adapt a read model or the views exposing the data
- When used with CQRS and eventual consistency is acceptable while the read model updates, or the performance impact of rehydrating entities from an event stream is acceptable
Sharding strategies
| Strategy | Trade-offs |
|---|---|
| Lookup | Offers more control over how shards are configured and used. Using virtual shards reduces the impact when rebalancing data, because new physical partitions can be added to even out the workload — and the mapping between a virtual shard and the physical partitions implementing it can be modified without affecting application code that uses a shard key. Looking up shard locations can impose additional overhead |
| Range | Easy to implement and works well with range queries, because they can often fetch multiple data items from a single shard in a single operation. Offers easier data management — if users in the same region are in the same shard, updates can be scheduled per time zone based on local load and demand. However it doesn't provide optimal balancing between shards; rebalancing is difficult and might not resolve uneven load if the majority of activity is for adjacent shard keys |
| Hash | Offers a better chance of even data and load distribution. Request routing can be accomplished directly using the hash function — there's no need to maintain a map. Computing the hash might impose additional overhead, and rebalancing shards is difficult |
Fault tolerance patterns
There are many fault-tolerant patterns covering all the stages of the error lifecycle; the course reviews a subset. The reference is Patterns for Fault Tolerant Software.
Three types of redundancy
| Type | Description |
|---|---|
| Spatial | The system has multiple copies in different places that are redundant, providing alternatives selectable with minimal unavailability. Easy to understand and visualize in terms of hardware — multiple copies of the hardware platform |
| Temporal | Redundancy that occurs over time, which helps achieve correct results but lengthens the time of unavailability. Recovery Blocks provide temporal software redundancy: a program consists of a primary block and secondary blocks; if the primary's result fails its acceptance test, the secondary blocks execute in sequence until a result passes. Recovery blocks offer primarily temporal redundancy because they execute sequentially |
| Informational | When information is repeated it can aid detection and correction. Provided by having multiple versions of the same data available, which can be stored in different places on different types of storage — for example disk and flash memory |
Spatial redundancy methods
| Method | Description |
|---|---|
| Active-Active | Provides totally redundant units of mitigation for the critical functionality. At any given time both elements are active and load sharing, and either is capable of processing the entire load. Provides the fastest recovery, but has a high cost because there is much more capability in the system than the ordinary workload needs. Implies a pairing between the two active units |
| Active-Standby | The standby is again paired with an active element, but the standby element is not performing useful work. This means the same amount of resources is needed as in active-active, but during normal operation the standby is idle |
| N+M | Where a one-to-one relationship is too expensive: N active elements process the workload and M redundant standby elements are ready to assume control when a failure occurs in any of the N |
The choice of redundancy regime has great effect on switching speed. If the redundant elements are hot standbys, switchover occurs very quickly with minimal outage of the main application. If the standby is only warm, it requires time to return to the same application state — checkpoints help get there more quickly. When the standby is cold, it must be started from an inactive state, which adds time; after starting it can be treated as a warm standby.
Related material covered: voting algorithms and triple modular redundancy.
Bulkhead
When to use: isolate resources used to consume a set of backend services, especially if the application can provide some level of functionality even when one of the services is not responding; isolate critical consumers from standard consumers; protect the application from cascading failures.
May not be suitable when: less efficient use of resources is not acceptable in the project, or the added complexity is not necessary.
Circuit breaker
Prevents an application from repeatedly trying to execute an operation that's likely to fail, allowing it to continue without waiting for the fault to be fixed or wasting CPU cycles determining that the fault is long-lasting. Fail fast!
It is a kind of proxy between the calling service and the called service, detecting whether the called service is in trouble — responses timing out or errors returned. When a threshold is exceeded the breaker changes state, preventing the caller from spending time and resources on calls unlikely to succeed anyway. The breaker must also be capable of detecting when the called service is up again, so it's fine to talk to it.
Use this pattern to prevent an application from trying to invoke a remote service or access a shared resource if the operation is highly likely to fail.
Two further fault-tolerance patterns are worth knowing by name. The Heartbeat pattern has a monitored component (or its monitor) emit a signal at regular intervals so failures are detected within a bounded time — the interval is a trade-off between faster detection and the monitoring overhead it adds. The Leaky Bucket Counter tracks how often a given error recurs over a decaying window — the bucket fills on each occurrence and drains at a steady rate — so occasional, transient errors are tolerated while an error rate that keeps exceeding the bucket's capacity is treated as permanent.
Security patterns
Federated identity
Using different credentials for multiple applications can:
- Cause a disjointed user experience — users often forget sign-in credentials when they have many different ones
- Expose security vulnerabilities — when a user leaves the company the account must immediately be deprovisioned, and it's easy to overlook this in large organizations
- Complicate user management — administrators must manage credentials for all users and perform additional tasks such as providing password reminders
When to use:
- Single sign-on in the enterprise — authenticate employees for corporate applications hosted in the cloud outside the corporate security boundary without requiring them to sign in every time. The experience matches on-premises applications, where they authenticate on signing in to the corporate network and thereafter have access to all relevant applications
- Federated identity with multiple partners — authenticate both corporate employees and business partners who don't have accounts in the corporate directory. Common in B2B applications, applications integrating with third-party services, and where companies with different IT systems have merged or shared resources
- Federated identity in SaaS applications — independent software vendors provide a ready-to-use service for multiple clients or tenants, each authenticating using a suitable identity provider. Business users use corporate credentials while consumers and clients of the tenant use social identity credentials
Might not be useful when: all users can be authenticated by one identity provider and there's no requirement to authenticate with another — typical in business applications using a corporate directory accessible within the application, via a VPN, or through a virtual network connection between the on-premises directory and the application. Also when the application was originally built with a different authentication mechanism, perhaps with custom user stores, or lacks the capability to handle the negotiation standards used by claims-based technologies — retrofitting claims-based authentication and access control into existing applications can be complex and probably not cost effective.
Gatekeeper
| Benefit | Detail |
|---|---|
| Controlled validation | The gatekeeper validates all requests and rejects those that don't meet validation requirements |
| Limited risk and exposure | The gatekeeper doesn't have access to the credentials or keys used by the trusted host to access storage and services. If the gatekeeper is compromised, the attacker doesn't get access to these credentials or keys |
| Appropriate security | The gatekeeper runs in a limited privilege mode while the rest of the application runs in the full trust mode required to access storage and services. If compromised, it can't directly access the application services or data |
Useful for: applications handling sensitive information, exposing services that must have a high degree of protection from malicious attacks, or performing mission-critical operations that shouldn't be disrupted; and distributed applications where it's necessary to perform request validation separately from the main tasks, or to centralize validation to simplify maintenance and administration.
Valet key
Useful when:
- To minimize resource loading and maximize performance and scalability — a valet key doesn't require the resource to be locked, no remote server call is required, there's no limit on the number of keys that can be issued, and it avoids a single point of failure resulting from performing the data transfer through application code. Creating a valet key is typically a simple cryptographic operation of signing a string with a key
- To minimize operational cost — enabling direct access to stores and queues is resource and cost efficient, can result in fewer network round trips, and might allow a reduction in the number of compute resources required
- When clients regularly upload or download data, particularly where there's a large volume or each operation involves large files
- When the application has limited compute resources available due to hosting limitations or cost. The pattern is even more helpful with many concurrent uploads or downloads, because it relieves the application from handling the transfer
- When data is stored in a remote data store or a different datacenter. If the application acted as a gatekeeper there might be a charge for the additional bandwidth of transferring data between datacenters, or across public or private networks between client, application and data store
Might not be useful when:
- The application must perform some task on the data before it's stored or sent — validation, logging access, or executing a transformation. However some data stores and clients can negotiate and carry out simple transformations such as compression and decompression (a web browser can usually handle GZip)
- The design of an existing application makes it difficult to incorporate the pattern — using it typically requires a different architectural approach for delivering and receiving data
- It's necessary to maintain audit trails or control the number of times a data transfer operation is executed, and the valet key mechanism in use doesn't support notifications the server can use to manage these operations
- It's necessary to limit the size of the data, especially during upload. The only solution is for the application to check the data size after the operation completes, or check the size of uploads after a specified period or on a scheduled basis
Zero-downtime deployment
It is essential to be able to roll back a deployment in case it goes wrong. Debugging problems in a running production environment is almost certain to result in late nights, mistakes with unfortunate consequences, and angry users. You need a way to restore service to your users when things go wrong, so you can debug the failure in the comfort of normal working hours.
Several rollback methods exist; the more advanced techniques — blue-green deployments and canary releasing — can also be used to perform zero-downtime releases and rollbacks.
Blue-green deployment. The idea is to have two identical versions of your production environment — blue and green. Users of the system are routed to the green environment, which is the currently designated production. In some situations these can be different pieces of hardware, or different virtual machines running on the same or different hardware; they can also be a single operating environment partitioned into separate zones with separate IP addresses for the two slices.
Canary release. It is usually a safe assumption that you only have one version of your software in production at a time. If you have an extremely large production environment it's impossible to create a meaningful capacity testing environment, unless your application's architecture employs end-to-end sharding. So how do you ensure a new version won't perform poorly?
Like blue-green, you initially deploy the new version to a set of servers where no users are routed. You can then do smoke tests and, if desired, capacity tests on the new version.
There are different strategies for choosing which users see the new version: a simple strategy is a random sample; some companies release to their internal users and employees before releasing to the world; a more sophisticated approach chooses users based on their profile and other demographics. You can even have multiple versions of your application in production at the same time, routing different groups of users to different versions as required.
Benefits:
- It makes rolling back easy — just stop routing users to the bad version and investigate the problem
- You can use it for A/B testing by routing some users to the new version and some to the old
- You can check whether the application meets capacity requirements by gradually ramping up the load, slowly routing more and more users while measuring response time and metrics like CPU usage, I/O and memory
Constraints:
- Hard to use where users have your software installed on their own computers or mobile devices
- Any shared resource needs to work with all versions of the application you want in production
- Supporting multiple versions is painful, so keep the number of canaries to a minimum
If your goal is zero downtime and your application has very low automated test coverage, canary release is the technique. In an extremely large production environment it's impossible to create a meaningful capacity testing environment, and a benefit of canary releases is the ability to do capacity testing of the new version in a production environment with a safe rollback strategy if issues are found. It lets you test on a small portion of users first, which is the least painful way to check functionality in the absence of auto-tests — though you still need a toolset for monitoring what's happening to users on the new version.
Integration
Enterprise Integration Patterns is the "bible" of the integration patterns — a must read. Concepts from the book are widely used in modern platforms, frameworks and tools.
Underneath that single reference sit four distinct integration styles — File Transfer, Shared Database, Remote Procedure Invocation, and Messaging — trading off coupling, latency and platform independence differently; messaging is the style most of the book's patterns assume. Individual patterns worth naming include the Messaging Gateway (encapsulates access to the messaging system behind a domain-specific API so the rest of the application never talks to the broker directly), Guaranteed Delivery (persists a message on both sender and receiver so it survives a broker crash), and the Aggregator (a stateful filter that collects a set of related messages and republishes them as one combined message once the set is complete).
Delivery model selection
The homework scenario: Lumen Diagnostics wants a scheduling platform, delivered as a service and hosted and supported by the Supplier. Which delivery model: IaaS, PaaS or SaaS?
SaaS. Lumen is buying a booking capability, not a place to build one. The supplier hosts and supports it; Lumen's staff configure rules and workflow rather than operating VMs. IaaS would dump operations onto Lumen. PaaS would mean Lumen (or each partner) still builds the application. The appendix already says "delivered as a service" — the interesting part of the answer is what that forces on multi-tenancy, data residency per region, and configuration vs custom code.
Exercises — Module 4
Task. Answer the questions in a short document. Try from your own knowledge first, then research. The layer figure is in the Layered section above. Two questions are set on Lumen Diagnostics.
Terms
- A design solution for a concrete context and problem: style, pattern, or tactic? How do the three differ?
- What are coupling and cohesion? Give a specific example of each.
Styles and patterns
- Typical cons of a monolith, including drawbacks you have seen
- The layer diagram in this module: is it a correct layered architecture, and why?
- Can C use B?
- Can A use C?
- When is event-driven worth using? When Mediator, when Broker?
- REST: is Richardson Level 3 worth chasing? Describe a case where it helps.
- What is a microservice? List challenges and drawbacks.
- When is CQRS a good fit?
- Name an integration framework or platform in a language you use; list pros and cons.
Resilience
- Error, Fault, Failure: which leads to which? One sample of each from a project you know.
- When does a circuit breaker help?
- Goal is zero downtime; the app has almost no automated tests. Which techniques, and why?
Lumen
- Lumen wants a supplier-hosted scheduling platform. IaaS, PaaS or SaaS?
- What does that choice force on tenancy, regional data, and configuration vs custom code?
- For Lumen, which style or pattern for each, and why against the case?
- (a) availability search
- (b) offer / confirmation workflow
- (c) one automated partner API plus ~600 manual partners