Skip to main content

PROD 110: APEX AIOps Incident Management Operator/End User Training

BASIC | 34 MIN

This course is for users who will be handling incidents in APEX AIOps Incident Management.  You will learn what Incident Management does and how it helps your day to day operations, then proceed to get familiar with the workflow of investigating and resolving incidents.

Whether you are completely new to Incident Management, or moving from Moogsoft Onprem, this course will get you ready to work with your own incidents.

BASIC | 4 MIN

With this video, get familiar with a typical workflow for a Moogsoft user.

Use case walkthrough: User workflow in APEX AIOps Incident Management

This video provides a use case walkthrough of what a typical workflow might look like for an Incident Management user as they resolve incidents.

In this video, we will step through the typical workflow of an APEX AIOps Incident Management user as they work through incidents.

Here comes a slack message, notifying us there’s a critical incident requiring our attention.

1_UW.jpg

So we click through to Incident Management, which takes us to this incident’s Situation Room.

The Situation Room is where you and your team can collaborate on an incident. This is the timeline for this incident. These sliders let you zoom in on particular areas, and the list below filters to match the time frame you choose.

2_UW.jpg

Let’s examine all alerts. These are the alerts that make up this incident.  Some of these are alerts from a monitoring system.

4_UW.jpg

And these are alerts generated by Incident Management based on the metrics it is tracking.

5_UW.jpg

We are going to own this incident.

6_UW.jpg

Now we will start our investigation.

These alerts came in within a few seconds of each other. Let’s look at the details of the alert that first came in.

7_UW.jpg

All attributes of this alert are visible now.

8_UW.jpg

And the metric information of the alert is visually presented here.

9_UW.jpg

Incident Management shows you the relevant context and their relationship to each other. This way, it’s much easier to grasp how the whole incident unfolded over time.

According to this, the volume queue length metric exceeded the threshold level and triggered a warning alert.

10_UW.jpg

Then the CPU usage metric on our front end server increased and triggered a warning alert...

11_UW.jpg

...the activity on the backend server fell...

12_UW.jpg

...and we are seeing a backend connection error critical alert.

13_UW.jpg

So, could this be the root cause that had a cascading effect to cause other alerts?

14_UW.jpg

Let’s check out the recommendations tab in the Situation Room for additional insight. Any similar incidents from the past will be surfaced here.

15_UW.jpg

Here's an incident that is 82% similar.

16_UW.jpg

And this icon means it has a resolving step we can review. Great!

17_UW.jpg

This indicates the similar incident was resolved using a runbook tool.

18_UW.jpg

Let’s go to this incident to get more context and confirm we can resolve our incident the same way.

19_UW.jpg

Let’s look at the comments. This incident involved a disk I/O bottleneck that showed up as an increase in Volume Queue Length. Just like our incident.

20_UW.jpg

We can use the same runbook tool to terminate runaway processes and free up resources.

21_UW.jpg

Let’s go back to our incident.

Currently the time window we are seeing is from the moment when the first alert in the incident occurred. We want to see what happens to the metrics when we run the tool. So let’s change the time frame. Now, the metrics are going to be updated in real time.

22_UW.jpg

We’ve run the tool, and with the runaway processes that were overloading I/O killed, the CPU load on the front-end web server is back to normal...

23_UW.jpg

...and activity has resumed on the back end server as well.

24_UW.jpg

Nice! The anomaly has resolved and now the metrics are within the normal range previously learned by the system. Good job!

25_UW.jpg

Now the alerts in our incident are all clear, as well as the incident itself.

26_UW.jpg

The incident status has been changed to resolved. We'll document our solution, and the case is closed!

27_UW.jpg

Just like that, we have resolved our first incident in Incident Management. Now it’s your turn to experience this workflow yourself!

Thanks for watching!

BASIC | 2 MIN

Learn how to manage your incident queue and dashboard.

Use case walkthrough: Queue and dashboards in APEX AIOps Incident Management

This video explains how you can use the queue and dashboards to view incidents in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

How do you know what’s coming down the pipeline for you to work on in Incident Management? This is the default incident list. This includes all open incidents, regardless of the nature of the issues or assignments.

Your operational procedure may be as simple as just looking at this list and picking an unassigned ticket.

queue_1.jpg

But most likely, you have a queue specific to your team. In this example, we have access to the Application Support team’s incidents view.

queue_2.jpg

So basically this is a list of application-related incidents. Your administrator may have set up a workflow to set the team assignment based on the impacted services, or there may be someone triaging incoming incidents and routing the applicable ones to your team.

queueextra.jpg

And you can make this your default view without affecting other users.

queue_3.jpg

Let’s say we are going to work on this one.

queue_4.jpg

There’s a list view that only shows the incidents you are assigned to.

queue_5.jpg

If you want to filter the list further you can do so here.

queue_extra_2.jpg
queue_6.jpg

You can save the view for yourself without affecting the original view.

queue_7.jpg

If you want to make it a shared view, you can do so here.

queue_9.jpg

When you create a view, you also get a corresponding dashboard.

queue_10.jpg

It presents data in a more visual manner, but you can also jump into a specific incident.

queue_extra_3.jpg

Now you know how to work with your incident queue. Thanks for watching!

BASIC | 2 MIN

In this video, you will learn how to create dashboards and share them in APEX AIOps Incident Management.

Use case walkthrough: Dashboards in APEX AIOps Incident Management

This video provides a use case walkthrough for using dashboards to easily view the performance of your teams and services in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

You can now create dashboards to view the team’s performance at a glance.

The incident list is useful for operations staff.

1_Dashboards_in_Moogsoft_Cloud_Edit.jpg

But as a manager you may need something to show how your teams or services are doing at a glance. The dashboard views are perfect for that. Here are the overall stats.

2_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Right now it’s tiled by service, but now it’s categorized by type.

3_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Or classes. You can slice and dice the overall category to suit your needs.

4_Dashboards_in_Moogsoft_Cloud_Edit.jpg

You can narrow down the list like this.

5_Dashboards_in_Moogsoft_Cloud_Edit.jpg

So if I wanted to learn more about critical application incidents it’s easy to do so. And, of course, go right into the Situation Room from here to start troubleshooting an incident.

6_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Once you get the exact data you are looking for, you can save that dashboard.

7_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Now you have one-click access to the team’s stats!

8_Dashboards_in_Moogsoft_Cloud_Edit.jpg

And the saved dashboards can be shared with specific groups or with everyone.

9_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Shared dashboards are accessible from here…

10_Dashboards_in_Moogsoft_Cloud_Edit.jpg

…and also from the Insight section.

11_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Now everyone can view this dashboard, and even set it as their default view!  Thanks for watching!

12_Dashboards_in_Moogsoft_Cloud_Edit.jpg

BASIC | 3 MIN

In this video, we will showcase the power of the Situation Room.

Use case walkthrough: Tour of the Situation Room ►

This video provides an overview of the Incident Management Situation Room and its features, which include the comments and recommendations tabs.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

In this video we’ll showcase the power of the Situation Room in Incident Management.

Here’s a critical incident happening…let’s take ownership of it and investigate.

1_Tour_of_Situation_Room_1.jpg

There is a lot going on–several Java Virtual Machines have crashed, and we’re seeing I/O and database problems.

2_Tour_of_Situation_Room_1.jpg

We’ll go to the Situation Room for this incident.

The Situation Room is a virtual collaboration space in Incident Management. It is designed to facilitate collaboration and drive incidents to resolution. Let me show you how it helps our investigation.

It has the same tools and information as the incident details page, but now the entire screen space is dedicated to resolving this one incident. You see the description of the incident, impacted services, and which correlation definition was applied to group the member alerts,

3_Tour_of_Situation_Room_1.jpg

But there are a few things that make the Situation Room special.

First, here is the comments tab. The team can chat as they work through the incident.

4_Tour_of_Situation_Room_1.jpg

Or maybe you just want to monitor the progress of this incident. Then you can add yourself, or a stakeholder as a watcher. You will receive an email when there are any announcements.

5_Tour_of_Situation_Room_1.jpg

Next, the recommendations tab is a great resource. Incident Management checks if there are incidents similar to the one at hand. And if there are, it will surface them for you to reference.

In this case, we have one similar incident.

6_Tour_of_Situation_Room_1.jpg

This one is 76% similar to our incident.

7_Tour_of_Situation_Room_1.jpg

And note this icon! This means information on what resolved this incident is available! With this past incident, it looks like the problems were related to a code push. Let’s learn more about it.

8_Tour_of_Situation_Room_1.jpg

Okay, now we have some more context.

9_Tour_of_Situation_Room_1.jpg

Let’s get back to the incident we were working on. Is there also a Jenkins alert in the current incident? Here’s the timeline that shows you how the incident unfolded…It says code deployment, so this is promising!

10_Tour_of_Situation_Room_1.jpg

Yes, that is the Jenkins alert. It’s likely that this code change is the root cause of the incident. So indeed, that similar incident Incident Management suggested was right!

11_Tour_of_Situation_Room_1.jpg

Let’s share our findings with the team. We’ll talk to the developers and get the deployment rolled back.

12_Tour_of_Situation_Room_1.jpg

All fixed… that was quick! Now you know how the Situation Room supports faster incident resolution. Thanks for watching!

13_Tour_of_Situation_Room_1.jpg

BASIC | 3 MIN

The Situation Room is where you can collaborate to share information and resolve incidents. In this video, learn about the top section of the Situation Room.

Use case walkthrough: Top Pane of Situation Room ►

This video explains how to use the different fields in the top pane of the Situation Room in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

Let’s take a closer look at each field of the situation room. We’ll focus on the top pane in this video.

1_Situation_Room_Top_Pane.jpg

This description of the incident is generated by the correlation definition that grouped the alerts.

2_Situation_Room_Top_Pane.jpg

It is defined here.

3_Situation_Room_Top_Pane.jpg

So in this example, the location, top three service names, number of sources, and top three source names are all dynamically inserted.

4_Situation_Room_Top_Pane.jpg

But you can edit it like this, if needed.

5_Situation_Room_Top_Pane.jpg

This shows the services impacted by this incident.

6_Situation_Room_Top_Pane.jpg

An incident can be assigned to an individual, and additionally, to one or more groups.

7_Situation_Room_Top_Pane.jpg

When it’s assigned to a person, the status changes.

8_Situation_Room_Top_Pane.jpg

How would you know when you have an incident assigned to you? A few ways. Here you can see all the incidents assigned to you.

9_Situation_Room_Top_Pane.jpg

Or your administrator may have configured an integration to trigger a notification.

10_Situation_Room_Top_Pane.jpg

The creation time is the time Incident Management grouped these alerts and created an incident. So note that it’s not the time the first event happened.

11_Situation_Room_Top_Pane.jpg

This shows how long the incident has been open.

12_Situation_Room_Top_Pane.jpg

If your administrator configured this, you can set a tag or perform tasks using a designated URL.

13_Situation_Room_Top_Pane.jpg

For example, in our environment you can go here to send this incident to ServiceNow.

14_Situation_Room_Top_Pane.jpg

Maybe you don’t need to be actively working on this incident, but just want to stay informed. Then click on the watch button.

15_Situation_Room_Top_Pane.jpg

You can add people other than yourself to watch the incident, too.

16_Situation_Room_Top_Pane.jpg

Now whenever there’s an announcement about this incident added here, the watchers will receive a notification.

17_Situation_Room_Top_Pane.jpg

If you want to set priority for your incidents, you can do so here. Then you can sort by priority and tackle the incidents with higher priority.

18_Situation_Room_Top_Pane.jpg

Now you are familiar with the top section of the situation room. Make sure to check out the other Situation Room deep dive videos!

BASIC | 2 MIN

This video shows you how to use the comments pane for collaboration, notifications, and keeping track of how you resolve incidents.

Use case walkthrough: Comments in Situation Room ►

This video steps through a use case for using comments in the Situation Room to collaborate with team members and stakeholders.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

The comments tab lets you collaborate with your team easily.

1_Comments_Tab_in_the_Situation_Room.jpg

As you work through an incident, all participants can chat in the comments tab.

2_Comments_Tab_in_the_Situation_Room.jpg

If you want to update stakeholders who aren’t actively working in the situation room with you, comment in the announcements tab.

3_Comments_Tab_in_the_Situation_Room.jpg

Your comment appears here like any other input...

4_Comments_Tab_in_the_Situation_Room.jpg

But it is also emailed to these people, keeping them informed about key progress. Someone who should be notified not listed here? You can add them!

5_Comments_Tab_in_the_Situation_Room.jpg

When you find out how to fix the problem, log that under the resolving steps tab.

6_Comments_Tab_in_the_Situation_Room.jpg

The input also shows up in the comment thread, but there’s more to it.

7_Comments_Tab_in_the_Situation_Room.jpg

Suppose a similar incident happens in future. Then Incident Management will suggest this incident as related...

8_Comments_Tab_in_the_Situation_Room.jpg

...with an indicator that there’s a resolving step.

9_Comments_Tab_in_the_Situation_Room.jpg

And the future 'you' will thank you for making it so easy to find how you fixed the problem last time.

10_Comments_Tab_in_the_Situation_Room.jpg

Alternatively, you can mark a regular comment as the resolving step after the fact.

11_Comments_Tab_in_the_Situation_Room.jpg

It works the same way as a comment you enter in the resolving steps tab.

Now you know how to use comments. Thanks for watching!

BASIC | 2 MIN

In this video, learn more about the recommendations tab in Situation Room.

Use case walkthrough: Recommendations tab in Situation Room ►

This video explains how to use the recommendations tab in Situation Room to reference similar incidents, suggest resolving steps, and expedite incident resolution.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

Let’s spend a few minutes learning about the recommendations tab in Situation Room.

1_Situation_Room_Recommendations.jpg

Incident Management checks if there are incidents similar to the one at hand. And if there are, it will surface them for you to reference.

2_Situation_Room_Recommendations.jpg

In this case, we have one similar incident.

3_Situation_Room_Recommendations.jpg

This one is 76% similar to our incident.

4_Situation_Room_Recommendations.jpg

And note this icon! This means information on what resolved this incident is available! With this past incident, it looks like the problems were related to a code push. Let’s learn more about it.

5_Situation_Room_Recommendations.jpg

Okay, now we have some more context.

6_Situation_Room_Recommendations.jpg

Now we can go back to the incident you are working on and see if it’s got that Jenkins alert.

And indeed, here it is. So just like this, the recommendations tab expedites your problem solving.

7_Situation_Room_Recommendations.jpg

But how did Incident Management surface that particular incident for us?

How Incident Management identifies similar incidents is configured here.

8_Situation_Room_Recommendations.jpg

By default, Incident Management compares these fields and tags to determine similarity. But you can change which fields to use.

10_Situation_Room_Recommendations.jpg

And how are resolving steps suggested?

It comes from comments that are tagged as resolving steps.

11_Situation_Room_Recommendations.jpg

So as you work on incidents, make sure to always mark the resolving steps. You will be glad you did in the future!

12_Situation_Room_Recommendations.jpg

BASIC | 1 MIN

Learn how to watch incidents in Moogsoft Cloud.

Use case walkthrough: Incident watcher in APEX AIOps Incident Management

This video explains how to watch incidents in APEX AIOps Incident Management and receive email notifications whenever announcements are added.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

There may be incidents you don’t need to directly work on, but just want to monitor progress. You can watch such incidents and stay informed.

1_Incident_Watcher.jpg

Now you are watching this incident. When anyone adds announcements, Incident Management will email you.

2_Incident_Watcher.jpg

Like this.

4_Incident_Watcher.jpg
3_Incident_Watcher.jpg

Note that only announcements trigger the notification emails. Also, you can add people other than yourself to incidents, like this:

5_Incident_Watcher.jpg

BASIC | 2 MIN

In this video, you will learn how to create incident list views and share them with others in Moogsoft Cloud.

Use case walkthrough: Shareable views in APEX AIOps Incident Management

This video provides a use case walkthrough on using sharable views in APEX AIOps Incident Management to customize the way the incidents list is displayed.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

Your incident list view is customizable. You can specify which columns are displayed and in what order.

1_Shareable_Views_in_Moogsoft_Cloud.jpg

And, you can save those custom views, and share them with your team! Let’s say you are in charge of the App Support group. You can create a filtered view like this and…

2_Shareable_Views_in_Moogsoft_Cloud.jpg

…save it.

3_Shareable_Views_in_Moogsoft_Cloud.jpg

Now you don’t have to reapply filters or rearrange the columns to see the team’s queue.

4_Shareable_Views_in_Moogsoft_Cloud.jpg

And you can share this view with your team.

5_Shareable_Views_in_Moogsoft_Cloud.jpg

Now everyone on your team is looking at exactly the same list, rendered in the exact same way. They can even save this view as their default view.

6_Shareable_Views_in_Moogsoft_Cloud.jpg

If you are managing multiple teams, you can set up a view for each team.

7_Shareable_Views_in_Moogsoft_Cloud.jpg

Also, you can switch to a dashboard view of the list. This is just an aggregated view of the same data. But with the dashboard view you can easily see the overall performance of the team, while still being able to drill down to a specific situation. For details about dashboards, watch this video.

9_Shareable_Views_in_Moogsoft_Cloud.jpg

Now you know how to create and share custom views with your team. Thanks for watching!

BASIC | 5 MIN

In this video, you’ll see how Moogsoft works to reduce operational noise.

Use case walkthrough: Deduplicate events to reduce noise ►

A busy service with multiple monitors can generate a flood of metrics, anomalies, and events. One issue might trigger a large number of repeat and duplicate events. APEX AIOps Incident Management analyzes every new piece of data — What is this? When did it happen? What is its severity? How often has it happened before? — and aggregates events for the same issue into alerts. Whenever it adds a new event, Incident Management updates the alert fields — event count, last event time, severity — so the alert always contains the latest information about the underlying issue. This process removes the duplicate, repeat, and obsolete noise from the data stream.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

One of the benefits of implementing Incident Management is that you can reduce noise, and focus on what matters. In this video, we’ll take a look at the noise reduction mechanism, using sample data from the real world.

We’ve extracted this sample data from an actual monitoring environment. The data is anonymized, but the events are real.

4_xls.png

The event data is coming from three different sources: Cloudwatch, a home grown monitoring tool, and Sensu. These three tools have been monitoring a SaaS applications environment.

5_diagram.png

Here we’ve got about 1700 events. Let’s feed this to Incident Management and see what happens.

Incident Management deduplicates events and forms alerts. For example, this event indicates disk i/o is very high for this.dpcc_p2 server. Then, here comes another event to tell you the condition has worsened.

9_2_events.png

These should not be considered separate problems, so Incident Management de-duplicates them into one alert.

10_Alert001.png

Let’s take a closer look. How exactly do we determine an event to be a duplicate?

In short, Incident Management compares the dedupe_key values of the two events, and if they are a 100% match, the two events are deduplicated into one alert. The dedupe_key value is a combination of the Source, Class, and Check fields in an event. If the incoming event is using the service field, that value is used also.

18_Reducing_Pager_Fatigue.jpg

In this example, instead of two events, you now have one alert in the critical state. Because of the deduplication, you will not be allocating separate resources to each event, and you will have more context for troubleshooting.

19_Reducing_Pager_Fatigue.jpg

Now consider this case: After the critical event, another event arrives with the same dedupe key value. But this time, the severity value of the event is Clear. Maybe an automated runbook kicked in and addressed the issue.

14_matching.png

Incident Management dedupes this event into an existing alert. The status of the alert updates to Clear.

15_Clear_alert.png

Without deduplication, it would take a manual correlation to figure out the issue has resolved itself. Have you had a series of link flapping events? All those up and down events would be consolidated into one alert with Incident Management.

So, after the 1700 events have been processed, here’s what we got. They are deduplicated into 29 alerts.

1_Reducing_Pager_Fatigue.jpg

This is the power of noise reduction. It helps you direct your focus on what actually matters.

But we don’t stop here. These 29 alerts are processed further before you come in. They are now correlated based on their relatedness.

Here, the 29 alerts are now clustered into 3 incidents.

2_Reducing_Pager_Fatigue.jpg

Incident Management evaluates alerts for their relatedness, and clusters the related alerts into one incident.To learn more about the correlation mechanism, watch the “Correlation Engine in APEX AIOps Incident Management” video.

3_Reducing_Pager_Fatigue.jpg

Imagine instead of getting paged 29 times for each of the alerts, now you get 3. No more pager fatigue for your team!

4_Reducing_Pager_Fatigue.jpg

Let’s examine one of the incidents and verify Incident Management has made a meaningful correlation. Let’s see if we can figure out what’s going on with this incident. Let’s go into the Situation Room. The Situation Room is where you can collaborate with your team on an incident.

5_Reducing_Pager_Fatigue.jpg

Looks like a team is already assigned to this one.

6_Reducing_Pager_Fatigue.jpg

Here, we can see the timeline of activities pertaining to this incident.

7_Reducing_Pager_Fatigue.jpg

We can zoom in to any particular area of interest to filter the activities.

8_Reducing_Pager_Fatigue.jpg

Looks like the group assigned is already reviewing this incident. Let’s do the same!

9_Reducing_Pager_Fatigue.jpg

We’ll look at the alerts in the incident. The incident has 10 alerts, which consist of 169 events.

11_Reducing_Pager_Fatigue.jpg

Lets see what else we can learn from the alert details. We can see that the dpcc system is the one involved.

12_Reducing_Pager_Fatigue.jpg

We have a monitoring setting to generate events when there is no activity for a prolonged amount of time. The hosts in this HA pair are not generating data...

13_Reducing_Pager_Fatigue.jpg

...and it looks like a core service is down.

14_Reducing_Pager_Fatigue.jpg

The internal message queues are backed up.

15_Reducing_Pager_Fatigue.jpg

Looking here, it looks like we are having problems with slow database writes. It’s possible that database problems might have caused the core processing service to fail.

16_Reducing_Pager_Fatigue.jpg

So, through Incident Management’s deduplication and correlation functionality, instead of looking at 1700 events to find 169 related events, you are presented with a complete picture from the start. Thanks for watching!

BASIC | 2 MIN

This one is another scenario to get you to work in Moogsoft in a guided manner.

Use case walkthrough: A tour of Incident Management for DevOps users ►

This video steps through an example of how a DevOps engineer might use Incident Management to resolve an incident.

Let’s sample a day in the life of a DevOps engineer in Incident Management. This incident just came in. We are going to assign it to us and investigate.

1_Dev.jpg

Let's go to the Situation Room for this incident. Judging from the description, the issue seems to be with the message writing service. Let’s take a look at the alerts that rolled up into this incident.

2_DevOps.jpg

Let’s take a look at the alerts that rolled up into this incident. This is the first alert. The description says we just deployed an updated message writer service.

3_DevOps.jpg

Then minutes later, the process duration metric went out of bounds to the critical state. So there seems to be a connection here.

4_Devops.jpg

Let’s verify this connection by looking at the metrics. By default, the metrics screen is set to show the incident range, but we want to see it in the context of what’s normal.

5_DevOps.jpg

So we are going to zoom out a bit… and set the time range to the last half hour.

6_DevOps.jpg

This is where the process duration metrics alert was triggered. Normally, it hardly takes 20 milliseconds to write anything to the database, and now it’s taking several seconds. That’s not good.

7_DevOps.jpg

Now we have a choice- we could rollback the deployment and fix the issue in code that’s adding seconds to this process, or accept this as a new norm and have Incident Management learn it.  For now, we are going to roll it back.

We rolled back the service update. And now a few minutes later, it looks like the metrics are back to normal. We’ll examine the message writing service and deploy an updated version when ready.

9_DevOps.jpg

We are going to go back to the incident and close it. We’ll describe what we did to resolve the incident here, that way, Incident Management can suggest a solution for similar incidents in the future.

10_DevOps.jpg

All set! Now you have seen a sample DevOps workflow in Incident Management. Thanks for watching!

11_DevOps.jpg

BASIC | 6 MIN

In this video, we will troubleshoot a sample incident to see the power of alert correlation.

Use case walkthrough: Power of alert correlation ►

This video steps through a use case example to showcase the power of alert correlation in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

We’ll troubleshoot a sample incident to see the power of alert correlation.

Here’s a simple setup that we’ll be using for our demonstration. We have a few different sources that are sending data to Incident Management.

26_Billing_Scenario.jpg

For starters, we’re ingesting metrics using the Metrics API. Incident Management is data agnostic, meaning that these metrics can be of any type, and can originate from any source.

We’re also using Splunk for logs and analytics, AppDynamics for application performance monitoring, and Prometheus for database and infrastructure performance monitoring.

And here’s what we are monitoring.

We have a three-tier, customer-facing application called Billing that relies on databases in Atlanta and New York.

Billing_Scenario_Service.jpeg

One of the switches in the Atlanta data center experiences some degradation - maybe some packet loss - which impedes the communication between the application server and database.

Billing_Scenario_Faulty_Switch

Since all three of these domains are monitored by different teams using different tools, we’ll see application alerts...

Billing_Scenario_Application_Alerts.jpeg

Database alerts...

Billing_Scenario_Database_Alerts.jpeg

And network alerts.

Billing_Scenario_Network_Alerts.jpeg

As an operator trying to resolve this problem, they often don’t have visibility into what other teams are seeing in their monitoring tools. So it takes time to synthesize your analysis of what you can see, with the insights from others.

That’s where Incident Management comes in. Incident Management identifies related alerts and groups them together, organizes them into relevant incidents, and presents these incidents with rich context, making it easy for us to identify the underlying problem.

Billing_Scenario_Correlation_Graphic.jpeg

Let me show you what I mean.

Here is the Incidents panel, where we’ll start our triage efforts.

And this is the incident you’d get for the scenario we just went over.

1_Billing_Scenario.jpg

Let's go into the Situation room to discover more. The service impacted by this incident is Billing...

2_Billing_Scenario.jpg

And the description has been dynamically populated to tell us the critical information - in this case, the locations affected by the incident.

3_Billing_Scenario.jpg

And this incident contains 11 alerts. No doubt, these came from all four data sources.

4_Billing_Scenario.jpg

Sure enough, we got some application related alerts from AppDynamics, database alerts via Prometheus and Splunk, infrastructure alerts from Prometheus, and network alerts from the Metrics API. Such context helps us decide which teams in our organization we might want to engage.

5_Billing_Scenario.jpg
6_Billing_Scenario.jpg

7_Billing_Scenario.jpg
8_Billing_Scenario.jpg

We are going to own this incident, and assign it to our group.

9_Billing_Scenario.jpg

Now let’s examine the alerts.

We’ll sort these alerts by First Event Time so we can see how the issue started, and how it evolved.

10_Billing_Scenario.jpg

You can see in the Manager column how this incident combines alerts from four different data sources.

11_Billing_Scenario.jpg

And under event count, you can see how Incident Management provides noise reduction by deduplicating events into alerts.

12_Billing_Scenario.jpg

Now let’s look at the individual alerts.

The very first alert came from a switch in the Atlanta data center. It looks like the packet drops went out of bounds.

14_Billing_Scenario.jpg

Only a few seconds after this alert was created, we began receiving additional alerts about database timeouts...

15_Billing_Scenario.jpg

Application page load failures...

16_Billing_Scenario.jpg

And various other issues.

17_Billing_Scenario.jpg

So from this, it seems that the packet drop alert might be the root cause behind this incident.

Let’s do some further investigation by looking at the metrics for this incident. We’ll only examine relevant data from the past hour.

18_Billing_Scenario.jpg

For each of these metrics, the pink line represents the raw values of the metric at any given interval. The light gray band in the background represents the normal operating range of the metric, as calculated by Incident Management.

19_Billing_Scenario.jpg

And these colored dots represent anomalies in the metric that Incident Management has detected and classified, in terms of significance.

20_Billing_Scenario.jpg

As the metric continued to stray out of bounds, the anomaly was classified as Critical...

21_Billing_Scenario.jpg

And was later cleared when the metric went back in bounds.

22_Billing_Scenario.jpg

This metric is the packet drops alert we talked about earlier. Let’s see if it’s really the underlying cause of this incident.

23_Billing_Scenario.jpg

If we compare this graph to the graphs of the other alerts, we can see that all the anomalies started as soon as packet drops were detected...

24_Billing_Scenario.jpg

And more importantly, all the anomalies were cleared as soon as packet drops were cleared.

25_Billing_Scenario.jpg

This is a pretty good indication that our hypothesis is correct. The packet drop issue seems to be causing all the other alerts.

With this figured out, we know exactly who we should talk to about next steps. We’ll notify the team responsible for maintaining our network, and ask them to check on the faulty switch in the Atlanta data center.

Imagine how much time we’ve saved just now by having all the info from separate source systems in one place, and having them grouped together based on their relatedness for you. Now you know how you can have Incident Management correlate alerts for you for a faster mean time to recover.

Thanks for watching!