← Coursework

Chiper Microservices Architecture Redesign

The architecture for a platform that supplies corner stores, taken from one application to services cut by business capability, kept in step by events and checked sprint by sprint against measured quality targets.

Course
Software Architecture
School
Universidad de Los Andes
Term
Fall 2019
Role
Software Architect
Stack
  • Django
  • Spring Boot
  • PostgreSQL
  • MongoDB
  • MQTT
  • Auth0
  • Amazon Web Services
  • Python
  • Java
  • Apache JMeter
  • SonarQube
Code
Before · sprints 1 to 3Owners and the CEOpages in a browserpagesLoad balancerin front of two copiesmaintenance page when degradedNginx · documented, no config in repo×2 copiesOne web applicationDjango 2.2 · 9 domain apps · one processstoresownerszonescatalogueordersorder linessalessale linesreportsone schemaOne relational databaseevery tablePostgreSQL on AWS RDSAfter · sprint 4Owners and the CEOthe same pagesShared databasecore and catalogue tablesPostgreSQL · canemdbpagesordersproductsCore · pages, ordersevery page, by rolecalls the other servicesDjango 2.2Catalogueowns the productspublishes each changeSpring Boot 2.1.3HTTPHTTPchangesReportszone suggestionsDjango · the core's projectEvent brokercatalogue changesMQTT · topic postzone reportseventsReport documentsone per zoneMongoDB AtlasStoresales, its catalogue copySpring Boot 2.1.3salesStore databaseits own tablesPostgreSQL · tiendadb
Before and after. Through three sprints the platform was one application on one database; the fourth sprint cut it into services by business capability, each reaching the data it needs.

Problem

A distributor that supplies thousands of small shops usually starts on one application that takes orders, holds the catalogue, records sales and reports on them. As the network grows, one heavy report or one busy week of orders slows everything down, and no part can change or scale without redeploying the whole. Moving such a platform to services is a series of choices rather than a rewrite: where to cut, which data each part owns, how parts that share data stay consistent, and how to prove each choice against a measured requirement. This project makes those choices for a supplier of corner stores in Latin America, one quality requirement per sprint.

Solution

The platform is cut by business capability into a core for orders and the owners' pages, a catalogue service, a store service for sales and a reports service, built in two technologies. The catalogue publishes every change to an event broker, and the store service keeps its own copy of the catalogue from those events, so it does not have to call the catalogue to work. Suggestions for each zone are ranked by a weekly job over the last seven days of sales and served as a stored report, so an owner's request reads one document instead of adding up millions of sale lines. Login and roles sit with an external identity provider, and each requirement was checked with a load test or a code scan against a seeded dataset of about seven million rows.

EnterCore web servicestore owner and CEO pages, by roleorder history, the last 5 orderscalls the other services over HTTPDjango 2.2 · 9 domain apps · EC2 :80Identity providerOAuth2 login, hosted outsiderole claim: store owner or CEOcredentials out of the core DBAuth0 · social-auth-app-django 3.1.0login · role claimServicesGET suggestionsGET /catalogolast 5 ordersReports servicezone suggestions for ownerssame Django project as the corereads a precomputed reportDjango · /recomendados · EC2 :8080Catalogue serviceowns the product cataloguecreate, read, update, deletepublishes every changeSpring Boot 2.1.3 · /catalogo · :9898Stores · eventszone reportcatalogue tablePOST · UPDATE · DELETEReport storeone document per zone reportread on every suggestionfilled through a POST endpointMongoDB Atlas · canemdb.reportesCore databasestores, orders, sales, zonesthe catalogue table, shared9 domain tablesPostgreSQL on AWS RDS · canemdbEvent brokerone topic for catalogue changespublish once, many subscribeat most once, last value keptMQTT :1883 · topic post · QoS 0Batch · subscriber7 days in · top 20 outcatalogue eventsWeekly ranking jobtop 20 products per zoneby sales over the last 7 daysweekly, off the request pathlogic_zonas.py · Saturday 22:00Store servicesales over RESTkeeps its own catalogue copysubscribes to catalogue eventsSpring Boot 2.1.3 · /ventas · :9898sales + replicaStore databasesales and store rowsits copy of the catalogueread by no other servicePostgreSQL on AWS RDS · tiendadb
Four services cut by business capability, in two technologies. The catalogue reaches the store service as events, not calls, and suggestions are ranked once a week, not on every request.

Learnings

  • Learning

    Each service owns its data

    The store service has its own database and learns about products only from events, so it can be deployed, scaled or restored without asking anyone. The core and the catalogue still wrote to one shared table, and that is exactly the coupling that makes two services move as one. A service is independent in production only when no other service reads or writes its tables.

  • Learning

    Events instead of calls between services

    Publishing a change once and letting any number of services subscribe removes the chain of synchronous calls, so a slow or stopped catalogue does not stop a sale. The price is that each copy is eventually right, not instantly right, and the delivery guarantee decides how far it can drift: at most once is cheap and can lose a change, while a copy that must stay correct needs at least once delivery and updates that are safe to apply twice. Choosing that guarantee on purpose is what makes replicated data trustworthy.

  • Learning

    Quality requirements as measured targets

    Every sprint turned a quality into a target and a test: how quickly an owner sees suggestions, how many order history requests the core serves each second. Ranking once a week instead of on every request is a choice that only makes sense against such a target, since it trades fresh rankings for an answer read from one stored document. Architecture decisions hold up in production only when each one is tied to a target that a test can pass or fail.

← Back to coursework