01
AI infrastructure deployment for training and inference
A data or engineering team has model experiments, growing inference demand, and no reliable way to allocate GPU capacity, protect datasets, or see who is using shared compute.
Typical flow: Identity and access → AI portal / API ingress → Kubernetes or workload scheduler → GPU compute pool → high-speed data and model storage → monitoring, quota, and audit services
When this becomes a project
GPU jobs compete with one another, teams copy datasets between machines, inference endpoints are managed manually, or a planned AI programme needs a defensible capacity and security design.
Delivery scope
- Profile training, batch, and inference workloads.
- Define tenancy, quota, and model release controls.
- Design the network, storage, and GPU scheduling model.
Indicative BoM / stack
- GPU servers sized after workload profiling.
- High-bandwidth fabric and resilient storage tiers.
- Kubernetes, model serving, registry, observability, and secrets management.
Implementation path
Start with a representative workload and a small compute pool. Test scheduling, data throughput, model deployment, access boundaries, and a node failure before adding capacity.
Acceptance checks
Documented workload sizing, repeatable environment provisioning, approved access model, monitored GPU use, and a test record for model deployment and recovery.
02
On-premises private cloud for shared enterprise workloads
An organisation has separate server islands, uneven utilisation, slow VM provisioning, and inconsistent backup or access controls. It wants a shared platform that stays in its own data centre.
Typical flow: User and admin access → identity services → cloud control plane → compute clusters → software-defined network → block, file, and object storage → backup and monitoring
When this becomes a project
New applications wait for infrastructure, hardware is purchased without a common capacity plan, and each team operates its own VM, network, and backup process.
Delivery scope
- Assess workloads, dependencies, and growth assumptions.
- Design tenancy, network zones, storage classes, and operational ownership.
- Plan pilot migration and service handover.
Indicative BoM / stack
- Compute nodes, redundant switching, and storage nodes.
- Virtualisation or cloud control plane with load balancing.
- Backup repository, monitoring, logging, and identity integration.
Implementation path
Build the management plane first, then add a workload zone. Migrate a low-risk application group before moving business-critical systems in agreed waves.
Acceptance checks
Tenant separation, tested backup and restore, VM provisioning workflow, capacity alerts, access review, and a documented procedure for routine operations.
03
Migration from on-premises systems to public cloud
A business wants to move selected applications out of its data centre, but the application inventory, network dependencies, data movement, and cutover ownership are still unclear.
Typical flow: Existing application and data estate → secure connectivity → cloud landing zone → application runtime and managed data services → backup, logging, and cost controls
When this becomes a project
Data-centre hardware reaches refresh date, teams need elastic capacity, or application owners need a safe way to move without treating every workload as a lift-and-shift candidate.
Delivery scope
- Inventory applications and map dependencies.
- Classify each workload: retain, rehost, replatform, refactor, or retire.
- Define connectivity, identity federation, migration waves, and rollback.
Indicative BoM / stack
- Cloud accounts, network hub, secure connectivity, and IAM roles.
- Compute, containers, managed database, and object storage where appropriate.
- Migration tooling, backup, logging, budget alerts, and security controls.
Implementation path
Establish the landing zone and connectivity before the first migration. Move one low-risk workload to validate the process, then schedule waves around dependency groups and business windows.
Acceptance checks
Approved application inventory, tested connectivity, verified data cutover, known rollback point, security and logging coverage, and a cost owner for each migrated service.
04
Consolidating legacy workloads into a private cloud
Servers and virtual machines have grown over time, but the organisation needs a standard platform for its existing estate without moving sensitive workloads to a public cloud.
Typical flow: Legacy servers and VMs → discovery and dependency map → private cloud landing zone → shared compute and storage → segmented network zones → backup, monitoring, and operations
When this becomes a project
Hardware renewal, inconsistent virtualisation practices, weak recovery processes, or a requirement to keep workloads and data within a controlled environment trigger the change.
Delivery scope
- Map the existing environment and group workloads into migration waves.
- Design the target platform and security zones.
- Define cutover, validation, and rollback for every wave.
Indicative BoM / stack
- Compute and storage clusters with redundant network paths.
- Private-cloud control plane, image catalogue, and automation tooling.
- Backup, replication, monitoring, and privileged-access controls.
Implementation path
Build and baseline the target platform, migrate a non-critical workload group, record the operating procedure, and then move subsequent waves after each acceptance review.
Acceptance checks
Successful application validation after migration, tested backup and restore, documented ownership, approved network segmentation, and a completed rollback test for the pilot wave.
05
Hybrid cloud setup with one operating model
An organisation needs some workloads to remain on premises while others use cloud services. The challenge is avoiding two disconnected platforms with duplicate identity, monitoring, and security processes.
Typical flow: On-premises applications and data → redundant private connectivity → cloud landing zone → shared identity and policy → central monitoring and logging → replicated data and backup services
When this becomes a project
Data Residency, latency, existing hardware investment, or specialised applications keep part of the estate local while new services need cloud elasticity or managed capabilities.
Delivery scope
- Define workload placement and integration boundaries.
- Design redundant connectivity, identity federation, and DNS patterns.
- Set common logging, backup, monitoring, and incident ownership.
Indicative BoM / stack
- Private connectivity with VPN failover where required.
- Federated identity, network firewalling, and central DNS.
- Cloud and on-premises monitoring, backup, and data replication services.
Implementation path
First establish identity and network connectivity. Then connect a single non-critical application flow, verify monitoring and support hand-offs, and scale the pattern to other workloads.
Acceptance checks
Tested primary and failover connectivity, single sign-on where planned, visible health signals across both environments, and a documented response path for cross-platform incidents.
06
Disaster Recovery as a Service for business-critical applications
Backup exists, but the business does not know how long recovery will take, which systems must be recovered together, or whether the recovery plan works outside a document.
Typical flow: Protected workloads → replication and backup policy → isolated recovery site or cloud environment → recovery orchestration → tested runbooks → monitoring, evidence, and review records
When this becomes a project
Audit requirements, outage exposure, ransomware risk, or a business decision to define realistic recovery objectives makes a repeatable recovery service necessary.
Delivery scope
- Run a business impact and dependency assessment.
- Set workload-specific RPO and RTO targets.
- Design replication, recovery order, communications, and test cadence.
Indicative BoM / stack
- Backup and replication tooling with immutable recovery copies.
- Recovery compute, network, storage, and secure administrative access.
- Runbooks, monitoring, test evidence, and service reporting.
Implementation path
Protect one application group first, complete a measured recovery exercise, resolve the gaps found, and then extend the service to the next business priority group.
Acceptance checks
Approved recovery objectives, successful restore and application validation, recovery-time evidence, named service owners, and an agreed schedule for future recovery tests.
07
Enterprise cloud foundation, governance, and landing zone
Cloud accounts are being created by separate teams. Access is inconsistent, logs are incomplete, and no one has a reliable view of policy compliance or ownership of spend.
Typical flow: Organisation hierarchy → account or subscription structure → identity federation → policy guardrails → network hub → central logs, security findings, and cost reporting
When this becomes a project
Cloud use has moved beyond a single team, audit questions are increasing, or a migration programme needs a stable foundation before the first production workload arrives.
Delivery scope
- Define account, subscription, and ownership model.
- Set identity, network, encryption, and logging baselines.
- Implement policy controls and an exception workflow.
Indicative BoM / stack
- Organisation management, identity federation, and privileged access.
- Network hub, DNS, firewalling, and private service access.
- Central logging, security posture management, and cost allocation tags.
Implementation path
Agree the operating rules with security, finance, and application owners first. Create a pilot account or subscription, validate the guardrails, and then use automation for every new environment.
Acceptance checks
Provisioning through a defined process, federated access, mandatory logging, policy violation reporting, cost allocation, and an approved exception route.
08
PaaS implementation and internal developer platform
Application teams spend too much time requesting environments, creating pipelines, and solving the same deployment problems. Security and operations teams then have to support each project differently.
Typical flow: Developer portal and templates → source control and CI → artifact registry → security checks → Kubernetes or application runtime → secrets, observability, and service catalogue
When this becomes a project
Teams need a faster but governed path from code to production, and the organisation wants standard templates rather than a new deployment method for every service.
Delivery scope
- Map the developer journey and current delivery bottlenecks.
- Design reusable service templates and golden paths.
- Integrate security checks, secrets, and operations signals into delivery.
Indicative BoM / stack
- Kubernetes or managed application runtime.
- Source control, CI/CD, artifact registry, and GitOps tooling.
- Secrets management, policy checks, observability, and developer portal.
Implementation path
Choose one application type and build its path end to end. Let a delivery team use it, fix the friction they find, then add other templates and a measured adoption plan.
Acceptance checks
A developer can provision an approved service path, deploy through controlled automation, view runtime health, rotate secrets, and recover a prior release without an ad hoc procedure.
09
VMware to KVM migration in controlled waves
A virtualisation estate is facing licensing, refresh, or operating-cost pressure. The aim is to migrate workloads without treating the hypervisor change as a simple file conversion.
Typical flow: VMware inventory and dependencies → target KVM or OpenStack cluster → network and storage mappings → test migration wave → production cutover → validation and decommissioning
When this becomes a project
Licensing change, hardware renewal, vendor strategy, or the need for an open virtualisation platform creates an exit case, but application compatibility and downtime risk remain unclear.
Delivery scope
- Inventory virtual machines, dependencies, and recovery needs.
- Design target compute, storage, network, and image standards.
- Plan migration waves, testing, rollback, and decommissioning.
Indicative BoM / stack
- KVM or OpenStack compute cluster and shared storage.
- Network segmentation, image management, and backup integration.
- Migration tooling, test environment, monitoring, and operating runbooks.
Implementation path
Build the target platform and prove it with a low-risk, representative workload. Only after compatibility and operational checks pass should the migration wave size increase.
Acceptance checks
Successful conversion and boot test, application-owner validation, backup coverage, mapped network policy, documented rollback, and a post-cutover performance review.
10
Private RAG and agent workflow for internal knowledge
Teams want faster answers from policies, documents, tickets, and internal systems, but cannot send all information to an uncontrolled external service or allow an AI workflow to act without review.
Typical flow: Identity-aware user access → application interface → retrieval and orchestration layer → approved document sources and tools → model endpoint → audit, evaluation, and human approval controls
When this becomes a project
Knowledge is spread across repositories, teams repeat the same research work, or the organisation needs an internal AI service with visibility into access, sources, and actions.
Delivery scope
- Classify source data and define access rules.
- Design retrieval, citations, evaluation, and model-hosting options.
- Set approval gates for tool use and a process for feedback and improvement.
Indicative BoM / stack
- Document connectors, ingestion pipeline, and vector or search store.
- Private or managed model endpoint with policy and identity gateway.
- Evaluation set, observability, audit logs, and workflow approval service.
Implementation path
Begin with one approved data source and a narrow question set. Evaluate answer quality and access enforcement with business owners before adding more sources or any action-taking tool.
Acceptance checks
Source-aware answers, access controls verified against user roles, documented evaluation results, audit records for tool calls, and clear human approval for sensitive actions.