Dynamic GRID extension to the cloud
A large Spanish financial group needed to absorb the risk-calculation peaks of its Treasury desk without adding hardware. We extended its GRID to the cloud with transparent autoscaling: −35% execution time, +200% capacity and −20% monthly compute cost.
Challenge
On-premise infrastructure saturated at the daily risk-calculation peaks (Basel and Bank of Spain requirements). Scaling by traditional means was slow and expensive and, during low-occupancy periods, the platform sat oversized.
Solution
Auto Scaling Groups that create and destroy cloud nodes according to the load detected by the calculation managers, integrated transparently into the existing GRID cluster: internal resources are used first and, if they are not enough, cloud nodes are activated and shut down when the work is done.
Technologies
- HPC / Grid Computing
- Cloud bursting with Auto Scaling Groups
- Hybrid on-premise + cloud architecture
- Automated node lifecycle
Context
The Treasury desk of a large Spanish financial group runs risk calculations that grow more complex and more frequent every year, driven by requirements such as Basel and the Bank of Spain's supervisory rules. Those calculations run on a corporate GRID that spreads tasks across thousands of on-premise cores.
The problem was not a lack of technology but an infrastructure that could not scale under pressure: daily peaks saturated the available capacity and, outside them, the platform sat oversized. The organisation needed a flexible, scalable solution that was compatible with the GRID already in place, without replacing it.
The challenge
Four points of friction concentrated the cost, the delays and the regulatory risk:
- Overload at critical moments: the daily calculation peaks, above all from Treasury, saturated the available infrastructure.
- Slow, expensive scaling: adding hardware was neither viable in the short term nor efficient in the long run.
- Low operational efficiency: during low-load periods the infrastructure was oversized, at a poorly optimised cost.
- Dependence on the local environment: the existing GRID architecture limited the adoption of more flexible options without proper integration.
Phased approach
The project ran in five phases, ensuring a progressive, controlled adoption aligned with the bank's operational requirements.
-
01
PoC and hybrid architecture
Initial technical validation and design of an architecture compatible with both the existing GRID and the cloud.
-
02
Autoscaling definition
Load metrics that activate nodes according to the real demand reported by the calculation managers.
-
03
Integration and testing
Validation of cloud nodes working as a transparent part of the compute cluster.
-
04
Automation and monitoring
Full orchestration of the node lifecycle and continuous supervision with preventive alerts.
-
05
Real workloads and tuning
Production workloads and fine-tuning of the scaling policies.
Architecture and strategy
The key was to extend, not to replace. Auto Scaling Groups create and destroy cloud nodes automatically according to the load detected by the calculation managers. Those nodes join the bank's compute cluster, so risk and treasury processes use them as if they were part of the on-premise environment.
The flow is simple: calculation tasks are managed from the central GRID; the fixed internal cores, which always run, are used first; if they are insufficient, additional cloud cores are activated automatically; when the tasks finish, those nodes shut themselves down. Full compatibility with the existing infrastructure, scaling adapted to real load, automated lifecycle of cloud resources and continuous monitoring with immediate reaction capacity.
Measurable results
The hybrid architecture turned an operational bottleneck into a competitive advantage: it did not just solve the capacity problem, it improved the performance of the risk and treasury systems.
| Indicator | Before | After |
|---|---|---|
| Total execution time of risk workloads | Saturated windows at peaks | −35% |
| Compute capacity | Limited to on-premise hardware | +200% with no physical investment |
| Monthly operating cost of compute | Permanent oversizing | −20% on average (pay-per-use) |
| Validated scalability | Manual hardware expansion | +1,500 simultaneous cores |
Aggregated figures, anonymised under a non-disclosure agreement. Source: Vermont Solutions projects.
Lessons that carry over
- Extend before you replace: you get the benefits of the cloud without disruption or risk to the Treasury operation.
- Autoscaling on demand removes oversizing and turns a fixed cost into a variable one.
- Transparent integration into the cluster avoids changing business processes: to the calculation manager, a cloud node is just another node.
- Continuous monitoring with preventive alerts is what makes automatic scaling trustworthy in a regulated environment.
Regulatory framework
A bank's risk calculation is subject to deadlines and supervisory requirements; compute capacity is therefore a compliance requirement, not just an efficiency one.
- Basel III/IV and the Bank of Spain's requirements: more frequent and more demanding capital and risk calculations.
- DORA (Art. 28): the cloud extension is designed as ICT third-party risk, with provider control and digital operational resilience.
- NIS2: network and information security in an essential sector; secure routing between the data centre and the cloud (load balancers, VPN, failover rules).
Frequently asked questions
Did the risk or treasury processes have to change to use the cloud?
No. Cloud nodes join the existing GRID cluster and the calculation managers treat them as their own cores. Business processes stay the same; what changes is the capacity behind them.
What happens to the cost when there are no load peaks?
Cloud nodes shut down on their own when the tasks finish. The pay-per-use model removes oversizing: the monthly operating cost of compute fell by 20% on average.
Does it apply to a GRID running on IBM Spectrum Symphony or TIBCO DataSynapse?
Yes. The design is independent of the calculation manager: the extension happens at cluster-node level. Vermont keeps a specialist team on both products and on migrating between them.
Can I know the institution and the full figures?
The case is anonymised under a non-disclosure agreement. Architecture details and full figures are shared after signing an NDA.
Related content
Last updated: 2026-09-12