Stock Groups

AWS explains outage and will make it easier to track future ones

[ad_1]

Amazon Web Services CEO Adam Selipsky gives a keynote speech at the AWS reInvent conference in Las Vegas, November 30, 2021.

Getty Images| Getty Images

AmazonWeb Services released Friday an explanation about the hours-long outage it experienced earlier this week. It included details regarding its retail operations and other online services. It also announced that the company plans to redesign its status page.

Amazon’s US East-1 Region of Virginia data centers experienced problems around 10:30 am. ET, Tuesday according to Amazon.

The company explained that an automated operation to increase the capacity of AWS service hosted on the main AWS network caused unexpected behavior by large numbers of clients within the network. a postOn its website. The result was that devices connected to AWS and Amazon’s internal networks became overwhelmed.

Several AWS services were damaged, including the widely-used EC2 service. This provides virtual server capacities. AWS engineers quickly resolved the issue and brought back service over the next few hours. EventBridge, which is a service that can be used to help developers create applications that respond to specific activities, did not fully recover until 9:40 PM. ET.

It can affect the impression that cloud infrastructure works well and is ready for migrations from physical data centers. This can have serious consequences for businesses. AWS serves millions of customers worldwide and has the leading providerIn the market

AWS has apologized to its customers for the disruption.

Popular sites and highly-used services, like Disney+ and Ticketmaster, were also taken offline. This outage affected Roomba vacuums, Amazon Ring security cameras, and other internet connected devices like smart cat litter box and ceiling fans that can be controlled by apps. 

Amazon’s U.S. retail operations came to an abrupt halt in certain areas. Amazon’s delivery and warehouse workforce use internal apps which rely upon AWS. This meant that employees could not scan packages, or get delivery routes. The site which manages orders for customers was not accessible to third party sellers.

AWS tried to notify customers of the situation during the outage but it was not possible for the cloud to update its information. status pageThis is the Service Health Dashboard.

AWS stated that the event had a singular cause. They decided to update customers via the Service Health Dashboard’s global banner. We have learned since that some customers find it hard to access information on this topic.

Customers were also unable to create support cases during the disruption for seven hours.

AWS indicated that they are now taking actions to fix both issues.

AWS announced that it will release an updated version of its Service Health Dashboard soon. It will be easier to see service impact and have a new support architecture that active runs in multiple AWS regions. We won’t experience delays communicating with our customers.

AWS is not the only one to make changes in the way that it reports problems.

2017 was a year that engineers couldn’t show the correct color for the Service Health Dashboard to indicate uptime due to an AWS S3 outage. Amazon created banners on Twitter and released new information.

Amazon announced that the SHD administration console has been modified to work across many AWS regions. a messageAbout that episode.

WATCH: The Week That Was: Amazon Web Services crash

[ad_2]