From Reactive AWS Maintenance to Proactive Cloud Governance
We stabilised a retailer’s AWS estate, modernised critical integrations and introduced an ongoing governance model for managing cloud risk, resilience, security and technical debt.
Our Customer’s Challenge
A supermarket retailer had developed a substantial AWS estate to support integrations between its operational systems, data platform and Shopify e-commerce environment. As the platform expanded, services and resources had accumulated across multiple AWS regions without a consistent process for identifying obsolescence, failed components, unnecessary resources or emerging operational risks.
An initial health assessment identified several immediate concerns. Numerous AWS Glue jobs were running on versions approaching end of support, legacy Amazon Pinpoint projects required decommissioning, and a backup plan had been failing daily because the file system it referenced no longer existed. Once these visible issues were addressed, the wider account also required a structured review to determine which regions and resources were genuinely in use.
The retailer needed more than a one-off technical clean-up. It required an ongoing service that could proactively monitor the AWS environment, investigate warnings, manage remediation work, challenge unnecessary complexity and provide clear governance over technical decisions and risks.
Our solution
We established a proactive AWS support and maintenance service combining estate discovery, technical remediation, root-cause analysis and continuous optimisation.
The first phase focused on stabilising the retailer’s critical integration landscape. We reviewed each legacy AWS Glue job, assessed its compatibility with supported versions and upgraded the jobs through a controlled programme of work. Where upgraded services continued to fail, we investigated the underlying cause rather than treating the upgrade itself as the assumed problem. This approach identified issues such as invalid database credentials that would otherwise have remained obscured by recurring job failures.
We also reviewed the retailer’s Amazon Pinpoint projects, confirmed their operational status and safely offboarded those that were no longer required. By October, all 16 identified Glue job upgrades and all five Pinpoint project offboarding activities had been completed.
The service then expanded into broader AWS governance and optimisation:
Auditing AWS resource usage region by region.
Separating genuinely active workloads from default or unused resources.
Assessing deactivation risks before making changes.
Restricting access to regions that could not be completely disabled.
Monitoring the environment following each change for unintended consequences.
Investigating newly discovered legacy services as the estate evolved.
Reviewing failed backup activity and removing obsolete configuration.
Producing a documented AWS backup strategy for future resilience planning.
Maintaining a prioritised backlog of security, reliability and cost-optimisation opportunities.
The regional review found several AWS regions containing only default supporting resources rather than active workloads. We recommended deactivating these regions to reduce unnecessary complexity, prevent resources from being created in the wrong location and limit the areas in which unauthorised activity could be concealed.
Following approval, the obsolete backup plan was deleted, a new backup strategy was delivered, two additional legacy Glue jobs were upgraded, and the first and second tranches of unused regions were deactivated or restricted through IAM controls. Further regional analysis was incorporated into the ongoing service backlog rather than treated as a one-off exercise.
The Results
- Integration modernisation: All 16 initially identified legacy AWS Glue jobs were upgraded, reducing exposure to unsupported technology and improving the maintainability of critical data integrations.
- Legacy service rationalisation: Five unused Amazon Pinpoint projects were safely offboarded, removing redundant services and future migration liabilities.
- Improved backup governance: A persistently failing backup configuration was removed and replaced with a documented backup strategy, giving the retailer a clearer foundation for agreeing future protection requirements.
- Reduced cloud complexity: Unused AWS regions were identified, assessed and progressively restricted, reducing the operational surface area of the account.
- Stronger security posture: Regional rationalisation reduced the opportunity for resources to be created in inappropriate locations or for unauthorised activity to remain unnoticed within rarely reviewed regions.
- Faster root-cause resolution: Recurring technical failures were investigated beyond their immediate symptoms, exposing underlying issues such as invalid credentials rather than relying on repeated restarts or superficial fixes.
- Controlled change delivery: Potentially disruptive changes were reviewed, approved and implemented in tranches, with post-change monitoring used to manage operational risk.
- Continuous improvement: The engagement established an ongoing cloud-governance capability that continually identifies technical debt, emerging risks, resilience gaps and optimisation opportunities.
More projects
Brand Update Delivery Management
Coordinating a complex cross-functional rebrand across digital, physical and operational channels; delivered on time and within budget.
Stock Knight: Turning Complex Financial Data into Confident Investment Decisions
We designed and developed Stock Knight, a scalable investment research platform that transforms complex financial data into accessible, decision-ready insights for equities and exchange-traded funds.