AI transformation across the infrastructure lifecycle: From supply chain to fleet operations

AI infrastructure is a system, and every part of that system is connected. Decisions made in silicon and systems design influence how infrastructure is sourced, deployed, and operated across the fleet. And what we learn once that hardware is running can inform what we build next.

This feedback matters because the system never stands still. Demand shifts, component constraints emerge, new capacity comes online, and hardware requirements evolve across the fleet. Keeping infrastructure reliable means continually learning and adapting as requirements change.

At Microsoft, Azure Hardware Systems and Infrastructure works across that lifecycle, from systems architecture and design through supply chain, deployment, and fleet operations across Azure’s more than 80 regions and 500 datacenter campuses. This end-to-end view gives us an opportunity to connect insights across the hardware lifecycle, so what we learn in one part of the system can improve decisions across the others.

Learn more about Azure infrastructure

AI can accelerate that learning. Across our own AI transformation, we are applying agentic and AI tools to help teams connect information, understand what is changing, and act sooner while keeping human judgment at the center. The opportunity is bigger than making individual tasks faster. It’s to build a system that learns from how infrastructure is designed, sourced, and operated, and applies those learnings to what comes next. This approach is part of our broader AI transformation journey.

Start with the work, not the AI

This process has reinforced a critical lesson as we’ve scaled how we apply AI as a force multiplier across our cloud infrastructure: AI transformation starts with the work, not the technology. Speed matters, but the greater opportunity is to redesign how decisions are made: what information is available when a decision needs to happen, how quickly teams can understand what changed, and where human judgment matters most.

Our cloud supply chain is a good example of this principle in practice. Every month, demand-planning teams forecast Azure’s infrastructure needs years into the future, accounting for changing customer demand, regional needs, installed capacity, and decommissioning activity. When a plan changes, determining why could require reconciling information across multiple systems, turning a single investigation into a lengthy process. It was tempting to look at that work and ask where we could add an agent.

For any company, putting AI on top of a fragmented process can simply make the fragmentation move faster. Before applying AI, our teams mapped and simplified the work, established a shared data foundation with quality, governance, and access controls, and identified decisions where people needed to remain accountable. We call this approach “Lean before AI.”

Starting with end-to-end processes and taking an AI-driven approach provides new ways of working and moves teams to parallel execution rather than sequential handoffs, resulting in integrated, collaborative workflows.

From days of research to decisions in minutes

With this foundation in place, we approached demand planning differently. A multi-agent workflow can examine signals such as installed-base shifts, regional demand, and decommissioning changes, then help planners understand what changed, where, and what drove the movement. Work that previously took five to seven business days can now be completed in hours, and sometimes in less than 20 minutes. Across more than five monthly planning cycles, our full demand-planning team saw approximately 50% less manual effort and cycle time fell by up to 75% in selected workflows.

The same pattern is taking shape across planning, product data, sourcing, fulfillment, logistics, and operational workflows. Specialized agents are helping teams spend less time finding and reconciling information and more time applying expertise.

In fulfillment, understanding why rack delivery is blocked from meeting customer demand could require teams to pull information manually from multiple sources. An intelligent assistant now brings together information about blockers and compatible or incompatible supplies, saving investigation time by as much as 55%. In logistics, an AI-powered logistics agent brings together data across air, land, and sea options to help teams evaluate speed, cost, as well as carbon tradeoffs and forecast emissions.

Building the learning loop

These individual applications matter, but the larger opportunity is to connect them. Our cloud supply chain team is moving toward end-to-end multi-agent workflows across bill-of-materials generation, capacity delivery, spare-parts management, capacity docking, and sales and operations execution. This work reflects a broader shift in our business: moving beyond isolated experiments toward a faster, more resilient, and intelligent operating system that places human judgment at the center.

The important outcome is not simply speed. Planners can begin with connected evidence instead of spending days assembling it, giving them more time to examine the explanation, add business context, and determine what it means for the decision ahead.

That is the learning loop we want. AI helps people reach the evidence faster. People bring context and judgment, act on what they learn, and create new information that can improve the next decision.

Learning across the fleet

The hardware lifecycle does not end when a server reaches a datacenter. Once infrastructure is deployed, the challenge becomes keeping it operating reliably for customers. Across millions of nodes in our fleet, continuous monitoring generates signals that help our teams investigate issues and determine root causes to decide how to act.

Across Azure, we’re moving cloud reliability upstream—transforming fleet management from reactive firefighting into a closed-loop system that prevents defects, predicts failures, and automatically restores hardware back into service. As we move toward a self-healing fleet, we are applying the same principles: connecting data across the lifecycle, continuous evaluation, redesigning workflows for human-agent orchestration, and keeping engineers in control of production decisions. Ultimately, this is also when the next learning cycle begins: systems collect and analyze information about the fleet, and failure patterns become evidence for how suppliers design and build the next generation of hardware.

Azure failure prediction and detection uses AI to analyze fleet telemetry so teams can identify emerging hardware failure patterns sooner and take action before potential issues impact customers. Engineers retain oversight of production decisions. This has already resulted in a 92% reduction in disk-related virtual machine (VM) interruptions and reduced repair time on out-of-service nodes by 53%. For rack managers, prediction provides up to three days of advance warning, enabling proactive recovery that cuts out-of-service repairs by 40%.

We also proactively and periodically screen our fleet to identify hardware vulnerable to silent data corruption before customer workloads are deployed, helping strengthen platform reliability.

We are also developing workflows that preserve context as hardware moves through investigation and recovery. For faulted resources, these workflows track assignment, action, outcome, and next step, with policy and approval controls around fleet actions. History and outcomes can then inform future decisions.

This is where the broader systems story comes together. Demand-planning decisions influence sourcing. Logistics affects when capacity reaches a datacenter. Once hardware is running, fleet telemetry and operational outcomes create another source of learning. Information about component performance can help teams proactively address potential issues before customers experience them. Those insights can also feed forward into the next generation of silicon, system, and rack design, while giving suppliers information to improve future components.

AI can make that loop faster, but the value comes from helping the system work better as a whole.

What we learned when things did not work

Some of our most useful lessons came from approaches that fell short.

We learned that applying AI to one part of a fragmented process can accelerate that task while creating more work somewhere else. An agent might produce its output faster, but if a downstream team must interpret, reformat or reconcile it manually, the workflow as a whole has not improved.

That changed how we measured success. Instead of evaluating only the task an agent performs, teams must examine the full workflow: the work removed, the new work created, the quality of the decision, and the outcome.

Reliable, accessible, and well-governed data is a prerequisite, not an afterthought. AI cannot compensate for conflicting definitions, unclear permissions, or information isolated across systems.

And we learned not to become attached to a particular architecture or agent. Models, frameworks, and business needs continue to change. A solution that is useful today may need to be redesigned six months from now or retired if the original business need no longer applies.

The resulting rhythm is practical: begin with a consequential decision, simplify the work around it, connect the right governed data, build alongside the people who know the work, evaluate the complete outcome, and keep changing as the business and technology evolve.

What we measure next

The next phase of enterprise AI will require us to measure more than adoption: how many people use an agent, how many agents are deployed, or how much time they save. Those measures matter, but they don’t tell us whether the work itself has improved. Can a planner understand in minutes a change that once took days to explain? Can a fulfillment manager resolve a capacity blocker without manually reconciling information across systems? Can an engineer identify a potential hardware problem before it becomes a customer problem? And can teams apply their expertise, so the next decision is better than the last?

That is the direction we are pursuing: a way of working that keeps learning as the technology and the business change. We are not trying to build the largest collection of agents or automate decisions simply because we can. We are working to produce more useful output from the infrastructure, information, and expertise already in the system.

The yield imperative

If AI is going to change what the world can build, we need to keep evolving how we build the infrastructure behind it. By connecting insights across silicon, systems, supply chain, and fleet operations, each stage can help improve the next. The result is reliable infrastructure ready when customers need it, and a system that learns how to deliver it better with every cycle.

Azure infrastructure

Learn more about Microsoft’s systems approach.

Get started

The post AI transformation across the infrastructure lifecycle: From supply chain to fleet operations appeared first on Microsoft Azure Blog.
Quelle: Azure

AWS Config now supports 77 new resource types

AWS Config now supports 77 additional AWS resource types across key services including Amazon EC2, Amazon S3 Files, and Amazon Q Business. This expansion provides greater coverage over your AWS environment, enabling you to more effectively discover, assess, audit, and remediate an even broader range of resources.
With this launch, if you have enabled recording for all resource types, then AWS Config will automatically track these new additions. The newly supported resource types are also available in Config rules and Config aggregators.
You can now use AWS Config to monitor the following newly supported resource types in all AWS Regions where the resources are available:

Resource Types:

AWS::ApplicationAutoScaling::ScalableTarget
AWS::GuardDuty::ThreatEntitySet
AWS::PCS::Queue

AWS::ApplicationSignals::Discovery
AWS::GuardDuty::TrustedEntitySet
AWS::QBusiness::DataSource

AWS::APS::AnomalyDetector
AWS::HealthLake::DataTransformationProfile
AWS::QBusiness::Index

AWS::APS::ResourcePolicy
AWS::InspectorV2::CodeSecurityScanConfiguration
AWS::QBusiness::Permission

AWS::ARCRegionSwitch::Plan
AWS::IoT::Logging
AWS::QBusiness::Plugin

AWS::Batch::QuotaShare
AWS::IoTWireless::WirelessDeviceImportTask
AWS::QBusiness::Retriever

AWS::Batch::ServiceEnvironment
AWS::LicenseManager::Grant
AWS::QBusiness::WebExperience

AWS::BedrockAgentCore::PaymentConnector
AWS::LicenseManager::License
AWS::QuickSight::OAuthClientApplication

AWS::BedrockAgentCore::ResourcePolicy
AWS::Lightsail::DatabaseSnapshot
AWS::RTBFabric::InboundExternalLink

AWS::Braket::SpendingLimit
AWS::MediaConnect::RouterInput
AWS::RTBFabric::RequesterGateway

AWS::CleanRooms::IdMappingTable
AWS::MediaConnect::RouterNetworkInterface
AWS::S3Files::AccessPoint

AWS::CleanRoomsML::ConfiguredModelAlgorithm
AWS::MediaLive::Multiplex
AWS::S3Files::FileSystem

AWS::CloudWatch::AlarmMuteRule
AWS::MediaLive::Node
AWS::S3Files::FileSystemPolicy

AWS::Connect::UserHierarchyStructure
AWS::MediaLive::SdiSource
AWS::S3Files::MountTarget

AWS::ConnectCampaignsV2::Campaign
AWS::MediaPackageV2::ChannelPolicy
AWS::SecurityAgent::Application

AWS::CUR::ReportDefinition
AWS::MediaPackageV2::OriginEndpointPolicy
AWS::SecurityAgent::TargetDomain

AWS::DataSync::LocationAzureBlob
AWS::MediaTailor::ChannelPolicy
AWS::SES::MailManagerAddressList

AWS::DataSync::LocationFSxONTAP
AWS::NeptuneGraph::GraphSnapshot
AWS::SSM::MaintenanceWindow

AWS::DataSync::LocationFSxOpenZFS
AWS::Notifications::EventRule
AWS::SSO::Application

AWS::EC2::TransitGatewayMeteringPolicy
AWS::Notifications::NotificationHub
AWS::Timestream::InfluxDBCluster

AWS::EC2::VPCCidrBlock
AWS::Omics::Configuration
AWS::Timestream::InfluxDBInstance

AWS::ECR::RegistryScanningConfiguration
AWS::Omics::WorkflowVersion
AWS::Transcribe::VocabularyFilter

AWS::EKS::Capability
AWS::OpenSearchServerless::CollectionGroup
AWS::Wisdom::AIGuardrail

AWS::ElementalInference::Feed
AWS::PCAConnectorAD::ServicePrincipalName
AWS::Wisdom::Assistant

AWS::EntityResolution::IdNamespace
AWS::PCS::Cluster
AWS::WorkSpacesWeb::NetworkSettings

AWS::Glue::Blueprint
AWS::PCS::ComputeNodeGroup
 

Quelle: aws.amazon.com

Claude Haiku 5.5 is now available on AWS GovCloud (US)

AWS GovCloud (US) now offers Claude Haiku 5.5, the fastest and most efficient model in the Claude 5.5 family, built for subagents and high-volume, cost-sensitive work. According to Anthropic, it costs around 75% less than Claude Haiku 4.5 for most tasks.
Claude Haiku 5.5 is Anthropic’s most capable Haiku yet, a significant step up from Haiku 4.5 across coding, tool use, computer use, and agents. It’s also the first Haiku with effort controls, so teams can tune cost against intelligence for each task. Haiku 5.5 fits work where speed and cost matter most: real-time experiences like voice agents, live support, and in-app assistants, and high-volume tasks like classification, summarization, and pulling key fields from documents. It’s also a fast, cost-efficient subagent. A more capable model like Claude Opus 5.5 can plan the work and hand off well-defined coding, tool use, and browser automation tasks to Haiku, making it practical to run many agents in parallel.
Amazon Bedrock gives you Haiku’s capabilities while keeping your data within AWS infrastructure with regional data residency and provides access through a unified service with AWS-managed features like Guardrails and Knowledge Bases. To learn more, see the Amazon Bedrock documentation and regional availability.
Quelle: aws.amazon.com

Claude Haiku 5.5 is now available on AWS

AWS now offers Claude Haiku 5.5, the fastest and most efficient model in the Claude 5.5 family, built for subagents and high-volume, cost-sensitive work. According to Anthropic, it costs around 75% less than Claude Haiku 4.5 for most tasks.
Claude Haiku 5.5 is Anthropic’s most capable Haiku yet, a significant step up from Haiku 4.5 across coding, tool use, computer use, and agents. It’s also the first Haiku with effort controls, so teams can tune cost against intelligence for each task. Haiku 5.5 fits work where speed and cost matter most: real-time experiences like voice agents, live support, and in-app assistants, and high-volume tasks like classification, summarization, and pulling key fields from documents. It’s also a fast, cost-efficient subagent. A more capable model like Claude Opus 5.5 can plan the work and hand off well-defined coding, tool use, and browser automation tasks to Haiku, making it practical to run many agents in parallel.
Customers have two ways to access Claude Haiku 5.5: Amazon Bedrock and Claude Platform on AWS.
Amazon Bedrock gives you Haiku’s capabilities while keeping your data within AWS infrastructure with regional data residency and provides access through a unified service with AWS-managed features like Guardrails and Knowledge Bases. To learn more, see the Amazon Bedrock documentation and regional availability.
Claude Platform on AWS gives you direct access to Anthropic’s native platform experience and capabilities via the AWS Console. Build, test, and deploy with the same APIs, features, and console experience you’d get working with Anthropic directly, unified with AWS billing and authentication. To get started, see the Claude Platform on AWS documentation.
Quelle: aws.amazon.com

AWS Capabilities by Region now offers availability notifications for individual features and advanced filters

Today, AWS announces feature-level availability notifications and advanced filters for AWS Capabilities by Region in AWS Builder Center. You can now set availability notifications for AWS service, or to an individual feature of an AWS service, in the AWS Regions you are interested in, and receive an in-app notification and a weekly email digest covering all the capabilities that became available since you enabled availability notifications. Setting availability notifications for a service covers every feature within it, and an ‘All regions’ option covers every Region, including AWS Regions that launch in the future.
In addition to the feature level availability notifications , you can now filter API operations and CloudFormation resources by availability status and compare Regions to show all capabilities, only matching ones, or only differences. With over 19,000 API operations and more than 5,000 CloudFormation resource types, these filters quickly narrow a full listing down to the exact differences between Regions, working alongside search and existing display preferences. 
Comparing capabilities in AWS Capabilities by Region requires no account. Setting up availability notifications requires a Builder ID, which is free. To get started, visit AWS Capabilities by Region on AWS Builder Center. To learn more, explore the AWS Builder Center FAQ
Quelle: aws.amazon.com