LOADING

Type to search

How Singapore GovTech Reduced the Cost of Generative AI by 75%

Cybersecurity Cybersecurity Studies & Reports

How Singapore GovTech Reduced the Cost of Generative AI by 75%

Share
the economics of generative AI

GovTech generative AI cost reduction shows how Singapore’s Government Technology Agency approached one of the practical challenges of deploying generative AI at scale: controlling costs without limiting the technology’s usefulness.

GovTech saw opportunities to use generative AI across public services, including streamlining processes and personalizing citizen interactions. But running large language models, or LLMs, at a national level presented a significant cost challenge. Its response was MAESTRO, a platform designed to provide cost-efficient, pre-built, and production-ready generative AI capabilities across government agencies.

According to the GovTech case study featured in the IBM Institute for Business Value’s Cybersecurity 2028: Your workforce, built for the AI frontier, GovTech optimized its computing infrastructure, selected high-performing foundation models, and used model quantization. It also used Amazon Bedrock and SageMaker JumpStart to select right-sized models and deploy multiple smaller, specialized models. The approach resulted in a 75% improvement in cost performance for generative AI workloads.

The cost problem behind enterprise generative AI

The economics of generative AI can change considerably when an organization moves from experimentation to regular use. A small pilot may involve limited users and relatively modest computing requirements. Supporting AI applications across multiple departments is a different proposition, particularly when those applications process large volumes of information.

For GovTech, the challenge was even broader because its objective was to support generative AI adoption across Singapore’s public sector. The organization needed an approach that could provide the required AI capabilities without making the cost of running large language models a barrier to wider adoption.

Rather than looking at the cost of an individual model in isolation, GovTech focused on the underlying infrastructure and the way different AI workloads were being handled. This included examining which models were appropriate for different requirements and how much computing capacity was actually needed.

That approach is relevant to organizations developing an enterprise generative AI strategy. As more AI applications move into production, the cost equation becomes less about the price of a single model and more about how efficiently the organization can operate its overall AI environment.

Choosing the right model for the workload

A central part of GovTech’s approach was the use of right-sized AI models. The case study explains that GovTech selected high-performing foundation models and used techniques such as model quantization. Amazon Bedrock and SageMaker JumpStart helped with model selection and deployment, allowing the organization to use multiple smaller, specialized models rather than relying on large, general-purpose models for every workload.

The result was lower resource consumption and lower costs while retaining the capabilities required for the applications. This is an important consideration for generative AI cost optimization, because the most powerful model available isn’t necessarily the most appropriate choice for every task.

Different workloads have different requirements. An application that performs a relatively focused task may not need the same model capacity as an application that requires more complex reasoning. Selecting a model based on the actual workload can therefore help organizations avoid paying for computing resources they don’t need.

This doesn’t mean larger models have no place in an enterprise AI environment. Some applications will require their capabilities. The lesson from GovTech is more practical: model selection should be driven by workload requirements rather than by the assumption that bigger is always better.

Where model quantization fits

Model quantization was another element of GovTech’s approach to AI model optimization. In broad terms, quantization reduces the numerical precision used within a model, which can reduce the memory and computing resources required to operate it.

The case study doesn’t present quantization as the sole reason for the 75% improvement in cost performance. It describes the technique as part of a broader strategy that combined model selection, computing infrastructure optimization and the use of smaller, specialized models.

That distinction is important when organizations assess AI workload optimization strategies. Cost efficiency rarely comes from one technical change. The model, infrastructure, workload and deployment approach all contribute to the resources required to operate an AI application.

For technology teams, this means optimization needs to be considered as part of the design process rather than treated as something to address after an AI application has already been deployed.

MAESTRO created a common foundation for AI

GovTech’s approach also addressed another issue that can affect the economics of AI adoption: duplicated infrastructure.

MAESTRO was developed as a platform to provide cost-efficient, pre-built and production-ready generative AI capabilities across government agencies. It included a no-code, unified, web-based interface for machine learning model building, training and deployment.

This helped make some machine learning capabilities accessible to people without deep technical expertise. It also provided a common environment in which teams could develop and deploy AI applications rather than having every project build its own underlying capabilities.

For organizations considering generative AI infrastructure optimization, this is worth considering. A shared platform won’t be appropriate for every organization or every workload, but where several teams have similar requirements, common infrastructure can reduce duplication and make technology management more consistent.

The broader point is that AI infrastructure needs to be planned around expected usage. An environment designed for a single experiment may not be suitable when AI becomes part of multiple business processes.

From cost optimization to operational value

Cost reduction is useful only when an organization continues to get value from the technology. GovTech’s case study provides examples of how the platform supported substantial workloads.

The Ministry of Manpower used MAESTRO to develop an AI “sensemaker” tool that processed more than one million documents in three months. The tool increased insight extraction by 60%, reduced sensemaking time by 50%, and saved more than 2,000 work hours.

The ministry also used the platform to develop an automated job classification tool. It processed 10 million job postings with 92% accuracy in three months.

These examples provide context for the cost-performance improvement. GovTech wasn’t optimizing its AI environment simply to reduce infrastructure spending. It was creating an environment in which government agencies could use AI for high-volume workloads and achieve measurable operational results.

That distinction is important for organizations evaluating the business case for generative AI. Cost should be considered alongside productivity, accuracy, processing volume and the value created by the application.

What this means for AI infrastructure planning

The GovTech case study points to a more workload-specific approach to AI infrastructure planning.

Organizations don’t necessarily need the same model, computing configuration or deployment approach for every AI application. A document-processing workload may have very different requirements from an application that performs complex reasoning or interacts with several enterprise systems.

Understanding those differences before deployment can help technology teams make better decisions about infrastructure and model selection. It also gives them a clearer basis for estimating the cost of operating AI as usage increases.

This is particularly important as organizations move from AI pilots to production. An application that works well with a small number of users may have very different infrastructure requirements when it becomes available to an entire business or government department.

Reducing generative AI costs therefore starts with understanding how the technology will actually be used. The model, workload and infrastructure need to be considered together.

Cost shouldn’t be separated from security

Cost optimization also needs to be considered alongside cybersecurity.

As organizations deploy more AI applications, security teams need visibility into the models being used, the data they process and the systems they can access. An AI environment that is cheaper to operate isn’t necessarily a better environment if it introduces unacceptable security or compliance risks.

This is particularly relevant for public-sector organizations that may process sensitive citizen information. Cost efficiency needs to coexist with appropriate controls around data, access and AI workloads.

A shared AI platform can help establish common approaches to deployment and management, but it also becomes part of the organization’s technology environment and needs to be secured accordingly.

For cybersecurity leaders, the practical objective is to balance cost, performance and security rather than optimize one at the expense of the others.

What organizations can learn from GovTech

The GovTech case study offers a practical framework for organizations looking at cost-effective generative AI adoption.

First, organizations need to understand the requirements of individual workloads before selecting a model. Not every application needs the capabilities of the largest model available.

Second, right-sized models can help reduce resource consumption where they provide sufficient performance. Model quantization can also be considered where it is appropriate for the deployment.

Third, organizations should look beyond the model itself when calculating AI costs. Computing infrastructure, deployment requirements and usage patterns all affect the total cost of running generative AI.

Finally, organizations with multiple AI teams should consider whether shared infrastructure can reduce duplication. A common platform can provide capabilities that individual teams would otherwise need to develop separately.

These decisions don’t guarantee a particular level of savings. They provide a structured way to approach AI cost management as deployments become larger and more complex.

The 75% figure is only part of the story

The most striking number in the GovTech case study is the 75% improvement in cost performance for generative AI workloads. But the figure becomes more meaningful when the decisions behind it are considered.

GovTech optimized its computing infrastructure, selected high-performing foundation models, used model quantization, and adopted right-sized models through Amazon Bedrock and SageMaker JumpStart. It also deployed multiple smaller, specialized models rather than relying on large general-purpose models for every workload.

The result was a reduction in resource consumption and cost while maintaining the capabilities required by the applications.

For organizations looking at GovTech generative AI cost reduction as an example for their own AI strategy, the important lesson isn’t to target the same percentage of savings. It’s to examine the choices that influence AI costs before the technology reaches production.

Generative AI will continue to require investment, but wider adoption doesn’t necessarily have to mean an equivalent increase in infrastructure costs. Organizations can assess their workloads, select appropriate models, optimize their computing environment and consider shared infrastructure where it makes sense.

GovTech’s experience shows how those decisions can work together. Its 75% improvement in cost performance was part of a broader approach to making generative AI practical at scale.

For technology and cybersecurity leaders, the takeaway is straightforward: the economics of generative AI are shaped by how models and infrastructure are selected, configured and deployed, not simply by the price of accessing an AI model.

Source

IBM Institute for Business Value, Cybersecurity 2028: Your workforce, built for the AI frontier. The report’s GovTech case study identifies the underlying source as an AWS case study on GovTech’s collaboration with the AWS Generative AI Innovation Center.

Author

  • New Project 18

    Maya Pillai is a technology writer with over 20 years of experience. She specializes in cybersecurity, focusing on ransomware, endpoint protection, and online threats, making complex issues easy to understand for businesses and individuals.

    View all posts
Tags:
Maya Pillai

Maya Pillai is a technology writer with over 20 years of experience. She specializes in cybersecurity, focusing on ransomware, endpoint protection, and online threats, making complex issues easy to understand for businesses and individuals.

  • 1