Skip to main content

Operations

How to use APEX AIOps Incident Management as an operator

BASIC | 2 MIN

Incident Management automatically closes old alerts and incidents behind the scenes. In this video, learn the default behavior for Auto-Close and where you can change the default settings.

Concept explainer: Auto-close in APEX AIOps Incident Management

This video explains how the auto-close feature works in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

Incident Management lists all open alerts and incidents, but let’s face it… you won’t investigate some of the alerts and incidents

1_Auto-Close.jpg

They did not lead to a major issue, or they are simply too old and you should rather be focusing on what’s impacting your world NOW.

3_1st_Explainer.png

Or, maybe you resolved an issue but just forgot to change the incident status to “closed”.

2_Auto-Close.jpg

To avoid more current and important issues from getting buried under older ones, Incident Management automatically closes the old alerts and incidents behind the scenes.

5_1st_Explainer.png

Here’s the default behavior. If all alerts in an incident are closed, then you no longer need to work on that incident.

6_2nd_Explainer.png

So Incident Management waits for 60 minutes, then auto-close the incident.

Auto-Close_2.png

Incidents are also auto-closed if it stays unclosed for 7 days. Because honestly, you won’t be working on such an old incident, would you?

9_3rd_Explainer.png

On the alert side, if an alert is resolved Incident Management will change the status of it after half an hour.  Alerts also gets closed if they are still open after 72 hours since it’s reported.

11_4th_explainer.png

You can change the default settings from here.

3_Auto-Close.jpg

But remember, incidents and alerts are interrelated. So not only a change in one incident can affect multiple alerts included in that incident,

14_final_screen.png

but also a change in status with one alert can affect multiple incidents.

15_final_screen.png

BASIC | 2 MIN

In this video, you will learn about the comments tab in the Situation Room.

Use case walkthrough: Comments in Situation Room ►

This video steps through a use case for using comments in the Situation Room to collaborate with team members and stakeholders.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

The comments tab lets you collaborate with your team easily.

1_Comments_Tab_in_the_Situation_Room.jpg

As you work through an incident, all participants can chat in the comments tab.

2_Comments_Tab_in_the_Situation_Room.jpg

If you want to update stakeholders who aren’t actively working in the situation room with you, comment in the announcements tab.

3_Comments_Tab_in_the_Situation_Room.jpg

Your comment appears here like any other input...

4_Comments_Tab_in_the_Situation_Room.jpg

But it is also emailed to these people, keeping them informed about key progress. Someone who should be notified not listed here? You can add them!

5_Comments_Tab_in_the_Situation_Room.jpg

When you find out how to fix the problem, log that under the resolving steps tab.

6_Comments_Tab_in_the_Situation_Room.jpg

The input also shows up in the comment thread, but there’s more to it.

7_Comments_Tab_in_the_Situation_Room.jpg

Suppose a similar incident happens in future. Then Incident Management will suggest this incident as related...

8_Comments_Tab_in_the_Situation_Room.jpg

...with an indicator that there’s a resolving step.

9_Comments_Tab_in_the_Situation_Room.jpg

And the future 'you' will thank you for making it so easy to find how you fixed the problem last time.

10_Comments_Tab_in_the_Situation_Room.jpg

Alternatively, you can mark a regular comment as the resolving step after the fact.

11_Comments_Tab_in_the_Situation_Room.jpg

It works the same way as a comment you enter in the resolving steps tab.

Now you know how to use comments. Thanks for watching!

BASIC | 2 MIN

Learn how to create dashboards in APEX AIOps Incident Management!

Use case walkthrough: Dashboards in APEX AIOps Incident Management

This video provides a use case walkthrough for using dashboards to easily view the performance of your teams and services in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

You can now create dashboards to view the team’s performance at a glance.

The incident list is useful for operations staff.

1_Dashboards_in_Moogsoft_Cloud_Edit.jpg

But as a manager you may need something to show how your teams or services are doing at a glance. The dashboard views are perfect for that. Here are the overall stats.

2_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Right now it’s tiled by service, but now it’s categorized by type.

3_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Or classes. You can slice and dice the overall category to suit your needs.

4_Dashboards_in_Moogsoft_Cloud_Edit.jpg

You can narrow down the list like this.

5_Dashboards_in_Moogsoft_Cloud_Edit.jpg

So if I wanted to learn more about critical application incidents it’s easy to do so. And, of course, go right into the Situation Room from here to start troubleshooting an incident.

6_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Once you get the exact data you are looking for, you can save that dashboard.

7_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Now you have one-click access to the team’s stats!

8_Dashboards_in_Moogsoft_Cloud_Edit.jpg

And the saved dashboards can be shared with specific groups or with everyone.

9_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Shared dashboards are accessible from here…

10_Dashboards_in_Moogsoft_Cloud_Edit.jpg

…and also from the Insight section.

11_Dashboards_in_Moogsoft_Cloud_Edit.jpg

Now everyone can view this dashboard, and even set it as their default view!  Thanks for watching!

12_Dashboards_in_Moogsoft_Cloud_Edit.jpg

BASIC | 1 MIN

In this video, learn how to use the incident watcher in the Situation Room.

Use case walkthrough: Incident watcher in APEX AIOps Incident Management

This video explains how to watch incidents in APEX AIOps Incident Management and receive email notifications whenever announcements are added.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

There may be incidents you don’t need to directly work on, but just want to monitor progress. You can watch such incidents and stay informed.

1_Incident_Watcher.jpg

Now you are watching this incident. When anyone adds announcements, Incident Management will email you.

2_Incident_Watcher.jpg

Like this.

4_Incident_Watcher.jpg
3_Incident_Watcher.jpg

Note that only announcements trigger the notification emails. Also, you can add people other than yourself to incidents, like this:

5_Incident_Watcher.jpg

BASIC | 2 MIN

As you work on incidents, sometimes you notice that you’ve seen the same problem before. You try the same solution, and the incident is resolved very quickly.   What if you don’t have to rely on your memory, but instead, have your Incident management system do this for you? In this video, you will learn about Probable Root Cause and how that can shorten the time to resolve.

Introduction to Probable Root Cause in APEX AIOps Incident Management

This video explains how to use the Probable Root Cause feature in APEX AIOps Incident Management to shorten the mean time to resolve.

As you work on incidents, sometimes you notice that you’ve seen the same problem before. You try the same solution, and the incident is resolved very quickly.

1_PRC.jpg

What if you don’t have to rely on your memory, but instead, have your Incident management system do this for you? That’s the idea of Probable Root Cause in Incident Management.

Make sure you mark the root cause alert every time you work on an incident.

2_PRC.jpg

Also, label other alerts as symptoms. You don't need to label every alert, but the more input you provide, the more Incident Management will learn.

3_PRC.jpg

Incident Management learns from your input. So the next time a similar incident occurs, it will recognize it.

4_PRC.jpg

And suggest the alert that most likely caused the incident.

5_PRC.jpg

So now, instead of inspecting all the alerts in this incident, you can zoom in on the critical alert immediately.

6_PRC.jpg

With Probable Root Cause, you can cut the time for analysis and shorten the time to resolve!

Here’s the documentation to learn more about the feature. Enjoy!Probable Root Cause overview

BASIC | 2 MIN

In this video, you will learn the different incident lists, how to create your own incidents, and how dashboards are created with created incident lists.

Use case walkthrough: Queue and dashboards in APEX AIOps Incident Management

This video explains how you can use the queue and dashboards to view incidents in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

How do you know what’s coming down the pipeline for you to work on in Incident Management? This is the default incident list. This includes all open incidents, regardless of the nature of the issues or assignments.

Your operational procedure may be as simple as just looking at this list and picking an unassigned ticket.

queue_1.jpg

But most likely, you have a queue specific to your team. In this example, we have access to the Application Support team’s incidents view.

queue_2.jpg

So basically this is a list of application-related incidents. Your administrator may have set up a workflow to set the team assignment based on the impacted services, or there may be someone triaging incoming incidents and routing the applicable ones to your team.

queueextra.jpg

And you can make this your default view without affecting other users.

queue_3.jpg

Let’s say we are going to work on this one.

queue_4.jpg

There’s a list view that only shows the incidents you are assigned to.

queue_5.jpg

If you want to filter the list further you can do so here.

queue_extra_2.jpg
queue_6.jpg

You can save the view for yourself without affecting the original view.

queue_7.jpg

If you want to make it a shared view, you can do so here.

queue_9.jpg

When you create a view, you also get a corresponding dashboard.

queue_10.jpg

It presents data in a more visual manner, but you can also jump into a specific incident.

queue_extra_3.jpg

Now you know how to work with your incident queue. Thanks for watching!

BASIC | 2 MIN

In this video, you will learn about the recommendations tab in the Situation Room.

Use case walkthrough: Recommendations tab in Situation Room ►

This video explains how to use the recommendations tab in Situation Room to reference similar incidents, suggest resolving steps, and expedite incident resolution.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

Let’s spend a few minutes learning about the recommendations tab in Situation Room.

1_Situation_Room_Recommendations.jpg

Incident Management checks if there are incidents similar to the one at hand. And if there are, it will surface them for you to reference.

2_Situation_Room_Recommendations.jpg

In this case, we have one similar incident.

3_Situation_Room_Recommendations.jpg

This one is 76% similar to our incident.

4_Situation_Room_Recommendations.jpg

And note this icon! This means information on what resolved this incident is available! With this past incident, it looks like the problems were related to a code push. Let’s learn more about it.

5_Situation_Room_Recommendations.jpg

Okay, now we have some more context.

6_Situation_Room_Recommendations.jpg

Now we can go back to the incident you are working on and see if it’s got that Jenkins alert.

And indeed, here it is. So just like this, the recommendations tab expedites your problem solving.

7_Situation_Room_Recommendations.jpg

But how did Incident Management surface that particular incident for us?

How Incident Management identifies similar incidents is configured here.

8_Situation_Room_Recommendations.jpg

By default, Incident Management compares these fields and tags to determine similarity. But you can change which fields to use.

10_Situation_Room_Recommendations.jpg

And how are resolving steps suggested?

It comes from comments that are tagged as resolving steps.

11_Situation_Room_Recommendations.jpg

So as you work on incidents, make sure to always mark the resolving steps. You will be glad you did in the future!

12_Situation_Room_Recommendations.jpg

BASIC | 2 MIN

Learn more about how to filter and create views on the Incidents details page.

Use case walkthrough: Shareable views in APEX AIOps Incident Management

This video provides a use case walkthrough on using sharable views in APEX AIOps Incident Management to customize the way the incidents list is displayed.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

Your incident list view is customizable. You can specify which columns are displayed and in what order.

1_Shareable_Views_in_Moogsoft_Cloud.jpg

And, you can save those custom views, and share them with your team! Let’s say you are in charge of the App Support group. You can create a filtered view like this and…

2_Shareable_Views_in_Moogsoft_Cloud.jpg

…save it.

3_Shareable_Views_in_Moogsoft_Cloud.jpg

Now you don’t have to reapply filters or rearrange the columns to see the team’s queue.

4_Shareable_Views_in_Moogsoft_Cloud.jpg

And you can share this view with your team.

5_Shareable_Views_in_Moogsoft_Cloud.jpg

Now everyone on your team is looking at exactly the same list, rendered in the exact same way. They can even save this view as their default view.

6_Shareable_Views_in_Moogsoft_Cloud.jpg

If you are managing multiple teams, you can set up a view for each team.

7_Shareable_Views_in_Moogsoft_Cloud.jpg

Also, you can switch to a dashboard view of the list. This is just an aggregated view of the same data. But with the dashboard view you can easily see the overall performance of the team, while still being able to drill down to a specific situation. For details about dashboards, watch this video.

9_Shareable_Views_in_Moogsoft_Cloud.jpg

Now you know how to create and share custom views with your team. Thanks for watching!

BASIC | 3 MIN

In this video, you will learn more about the top pane in the Situation Room.

Use case walkthrough: Top Pane of Situation Room ►

This video explains how to use the different fields in the top pane of the Situation Room in APEX AIOps Incident Management.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

Let’s take a closer look at each field of the situation room. We’ll focus on the top pane in this video.

1_Situation_Room_Top_Pane.jpg

This description of the incident is generated by the correlation definition that grouped the alerts.

2_Situation_Room_Top_Pane.jpg

It is defined here.

3_Situation_Room_Top_Pane.jpg

So in this example, the location, top three service names, number of sources, and top three source names are all dynamically inserted.

4_Situation_Room_Top_Pane.jpg

But you can edit it like this, if needed.

5_Situation_Room_Top_Pane.jpg

This shows the services impacted by this incident.

6_Situation_Room_Top_Pane.jpg

An incident can be assigned to an individual, and additionally, to one or more groups.

7_Situation_Room_Top_Pane.jpg

When it’s assigned to a person, the status changes.

8_Situation_Room_Top_Pane.jpg

How would you know when you have an incident assigned to you? A few ways. Here you can see all the incidents assigned to you.

9_Situation_Room_Top_Pane.jpg

Or your administrator may have configured an integration to trigger a notification.

10_Situation_Room_Top_Pane.jpg

The creation time is the time Incident Management grouped these alerts and created an incident. So note that it’s not the time the first event happened.

11_Situation_Room_Top_Pane.jpg

This shows how long the incident has been open.

12_Situation_Room_Top_Pane.jpg

If your administrator configured this, you can set a tag or perform tasks using a designated URL.

13_Situation_Room_Top_Pane.jpg

For example, in our environment you can go here to send this incident to ServiceNow.

14_Situation_Room_Top_Pane.jpg

Maybe you don’t need to be actively working on this incident, but just want to stay informed. Then click on the watch button.

15_Situation_Room_Top_Pane.jpg

You can add people other than yourself to watch the incident, too.

16_Situation_Room_Top_Pane.jpg

Now whenever there’s an announcement about this incident added here, the watchers will receive a notification.

17_Situation_Room_Top_Pane.jpg

If you want to set priority for your incidents, you can do so here. Then you can sort by priority and tackle the incidents with higher priority.

18_Situation_Room_Top_Pane.jpg

Now you are familiar with the top section of the situation room. Make sure to check out the other Situation Room deep dive videos!

BASIC | 3 MIN

Take a tour of the Situation Room in APEX AIOps Incident Management!

Use case walkthrough: Tour of the Situation Room ►

This video provides an overview of the Incident Management Situation Room and its features, which include the comments and recommendations tabs.

*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.

In this video we’ll showcase the power of the Situation Room in Incident Management.

Here’s a critical incident happening…let’s take ownership of it and investigate.

1_Tour_of_Situation_Room_1.jpg

There is a lot going on–several Java Virtual Machines have crashed, and we’re seeing I/O and database problems.

2_Tour_of_Situation_Room_1.jpg

We’ll go to the Situation Room for this incident.

The Situation Room is a virtual collaboration space in Incident Management. It is designed to facilitate collaboration and drive incidents to resolution. Let me show you how it helps our investigation.

It has the same tools and information as the incident details page, but now the entire screen space is dedicated to resolving this one incident. You see the description of the incident, impacted services, and which correlation definition was applied to group the member alerts,

3_Tour_of_Situation_Room_1.jpg

But there are a few things that make the Situation Room special.

First, here is the comments tab. The team can chat as they work through the incident.

4_Tour_of_Situation_Room_1.jpg

Or maybe you just want to monitor the progress of this incident. Then you can add yourself, or a stakeholder as a watcher. You will receive an email when there are any announcements.

5_Tour_of_Situation_Room_1.jpg

Next, the recommendations tab is a great resource. Incident Management checks if there are incidents similar to the one at hand. And if there are, it will surface them for you to reference.

In this case, we have one similar incident.

6_Tour_of_Situation_Room_1.jpg

This one is 76% similar to our incident.

7_Tour_of_Situation_Room_1.jpg

And note this icon! This means information on what resolved this incident is available! With this past incident, it looks like the problems were related to a code push. Let’s learn more about it.

8_Tour_of_Situation_Room_1.jpg

Okay, now we have some more context.

9_Tour_of_Situation_Room_1.jpg

Let’s get back to the incident we were working on. Is there also a Jenkins alert in the current incident? Here’s the timeline that shows you how the incident unfolded…It says code deployment, so this is promising!

10_Tour_of_Situation_Room_1.jpg

Yes, that is the Jenkins alert. It’s likely that this code change is the root cause of the incident. So indeed, that similar incident Incident Management suggested was right!

11_Tour_of_Situation_Room_1.jpg

Let’s share our findings with the team. We’ll talk to the developers and get the deployment rolled back.

12_Tour_of_Situation_Room_1.jpg

All fixed… that was quick! Now you know how the Situation Room supports faster incident resolution. Thanks for watching!

13_Tour_of_Situation_Room_1.jpg

BASIC | 4 MIN

Learn how a user might work in Incident Management. After watching this video, you will be able to identify the typical workflow of a user.

Use case walkthrough: User workflow in APEX AIOps Incident Management

This video provides a use case walkthrough of what a typical workflow might look like for an Incident Management user as they resolve incidents.

In this video, we will step through the typical workflow of an APEX AIOps Incident Management user as they work through incidents.

Here comes a slack message, notifying us there’s a critical incident requiring our attention.

1_UW.jpg

So we click through to Incident Management, which takes us to this incident’s Situation Room.

The Situation Room is where you and your team can collaborate on an incident. This is the timeline for this incident. These sliders let you zoom in on particular areas, and the list below filters to match the time frame you choose.

2_UW.jpg

Let’s examine all alerts. These are the alerts that make up this incident.  Some of these are alerts from a monitoring system.

4_UW.jpg

And these are alerts generated by Incident Management based on the metrics it is tracking.

5_UW.jpg

We are going to own this incident.

6_UW.jpg

Now we will start our investigation.

These alerts came in within a few seconds of each other. Let’s look at the details of the alert that first came in.

7_UW.jpg

All attributes of this alert are visible now.

8_UW.jpg

And the metric information of the alert is visually presented here.

9_UW.jpg

Incident Management shows you the relevant context and their relationship to each other. This way, it’s much easier to grasp how the whole incident unfolded over time.

According to this, the volume queue length metric exceeded the threshold level and triggered a warning alert.

10_UW.jpg

Then the CPU usage metric on our front end server increased and triggered a warning alert...

11_UW.jpg

...the activity on the backend server fell...

12_UW.jpg

...and we are seeing a backend connection error critical alert.

13_UW.jpg

So, could this be the root cause that had a cascading effect to cause other alerts?

14_UW.jpg

Let’s check out the recommendations tab in the Situation Room for additional insight. Any similar incidents from the past will be surfaced here.

15_UW.jpg

Here's an incident that is 82% similar.

16_UW.jpg

And this icon means it has a resolving step we can review. Great!

17_UW.jpg

This indicates the similar incident was resolved using a runbook tool.

18_UW.jpg

Let’s go to this incident to get more context and confirm we can resolve our incident the same way.

19_UW.jpg

Let’s look at the comments. This incident involved a disk I/O bottleneck that showed up as an increase in Volume Queue Length. Just like our incident.

20_UW.jpg

We can use the same runbook tool to terminate runaway processes and free up resources.

21_UW.jpg

Let’s go back to our incident.

Currently the time window we are seeing is from the moment when the first alert in the incident occurred. We want to see what happens to the metrics when we run the tool. So let’s change the time frame. Now, the metrics are going to be updated in real time.

22_UW.jpg

We’ve run the tool, and with the runaway processes that were overloading I/O killed, the CPU load on the front-end web server is back to normal...

23_UW.jpg

...and activity has resumed on the back end server as well.

24_UW.jpg

Nice! The anomaly has resolved and now the metrics are within the normal range previously learned by the system. Good job!

25_UW.jpg

Now the alerts in our incident are all clear, as well as the incident itself.

26_UW.jpg

The incident status has been changed to resolved. We'll document our solution, and the case is closed!

27_UW.jpg

Just like that, we have resolved our first incident in Incident Management. Now it’s your turn to experience this workflow yourself!

Thanks for watching!