Operations
How to use APEX AIOps Incident Management as an operator
BASIC | 2 MIN
Incident Management automatically closes old alerts and incidents behind the scenes. In this video, learn the default behavior for Auto-Close and where you can change the default settings.
Concept explainer: Auto-close in APEX AIOps Incident Management ►
This video explains how the auto-close feature works in APEX AIOps Incident Management.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
Incident Management lists all open alerts and incidents, but let’s face it… you won’t investigate some of the alerts and incidents

They did not lead to a major issue, or they are simply too old and you should rather be focusing on what’s impacting your world NOW.

Or, maybe you resolved an issue but just forgot to change the incident status to “closed”.

To avoid more current and important issues from getting buried under older ones, Incident Management automatically closes the old alerts and incidents behind the scenes.

Here’s the default behavior. If all alerts in an incident are closed, then you no longer need to work on that incident.

So Incident Management waits for 60 minutes, then auto-close the incident.

Incidents are also auto-closed if it stays unclosed for 7 days. Because honestly, you won’t be working on such an old incident, would you?

On the alert side, if an alert is resolved Incident Management will change the status of it after half an hour. Alerts also gets closed if they are still open after 72 hours since it’s reported.

You can change the default settings from here.

But remember, incidents and alerts are interrelated. So not only a change in one incident can affect multiple alerts included in that incident,

but also a change in status with one alert can affect multiple incidents.

BASIC | 2 MIN
In this video, you will learn about the comments tab in the Situation Room.
Use case walkthrough: Comments in Situation Room ►
This video steps through a use case for using comments in the Situation Room to collaborate with team members and stakeholders.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
The comments tab lets you collaborate with your team easily.

As you work through an incident, all participants can chat in the comments tab.

If you want to update stakeholders who aren’t actively working in the situation room with you, comment in the announcements tab.

Your comment appears here like any other input...

But it is also emailed to these people, keeping them informed about key progress. Someone who should be notified not listed here? You can add them!

When you find out how to fix the problem, log that under the resolving steps tab.

The input also shows up in the comment thread, but there’s more to it.

Suppose a similar incident happens in future. Then Incident Management will suggest this incident as related...

...with an indicator that there’s a resolving step.

And the future 'you' will thank you for making it so easy to find how you fixed the problem last time.

Alternatively, you can mark a regular comment as the resolving step after the fact.

It works the same way as a comment you enter in the resolving steps tab.
Now you know how to use comments. Thanks for watching!
BASIC | 2 MIN
Learn how to create dashboards in APEX AIOps Incident Management!
Use case walkthrough: Dashboards in APEX AIOps Incident Management ►
This video provides a use case walkthrough for using dashboards to easily view the performance of your teams and services in APEX AIOps Incident Management.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
You can now create dashboards to view the team’s performance at a glance.
The incident list is useful for operations staff.

But as a manager you may need something to show how your teams or services are doing at a glance. The dashboard views are perfect for that. Here are the overall stats.

Right now it’s tiled by service, but now it’s categorized by type.

Or classes. You can slice and dice the overall category to suit your needs.

You can narrow down the list like this.

So if I wanted to learn more about critical application incidents it’s easy to do so. And, of course, go right into the Situation Room from here to start troubleshooting an incident.

Once you get the exact data you are looking for, you can save that dashboard.

Now you have one-click access to the team’s stats!

And the saved dashboards can be shared with specific groups or with everyone.

Shared dashboards are accessible from here…

…and also from the Insight section.

Now everyone can view this dashboard, and even set it as their default view! Thanks for watching!

BASIC | 1 MIN
In this video, learn how to use the incident watcher in the Situation Room.
Use case walkthrough: Incident watcher in APEX AIOps Incident Management ►
This video explains how to watch incidents in APEX AIOps Incident Management and receive email notifications whenever announcements are added.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
There may be incidents you don’t need to directly work on, but just want to monitor progress. You can watch such incidents and stay informed.

Now you are watching this incident. When anyone adds announcements, Incident Management will email you.

Like this.


Note that only announcements trigger the notification emails. Also, you can add people other than yourself to incidents, like this:

BASIC | 2 MIN
As you work on incidents, sometimes you notice that you’ve seen the same problem before. You try the same solution, and the incident is resolved very quickly. What if you don’t have to rely on your memory, but instead, have your Incident management system do this for you? In this video, you will learn about Probable Root Cause and how that can shorten the time to resolve.
Introduction to Probable Root Cause in APEX AIOps Incident Management ►
This video explains how to use the Probable Root Cause feature in APEX AIOps Incident Management to shorten the mean time to resolve.
As you work on incidents, sometimes you notice that you’ve seen the same problem before. You try the same solution, and the incident is resolved very quickly.

What if you don’t have to rely on your memory, but instead, have your Incident management system do this for you? That’s the idea of Probable Root Cause in Incident Management.
Make sure you mark the root cause alert every time you work on an incident.

Also, label other alerts as symptoms. You don't need to label every alert, but the more input you provide, the more Incident Management will learn.

Incident Management learns from your input. So the next time a similar incident occurs, it will recognize it.

And suggest the alert that most likely caused the incident.

So now, instead of inspecting all the alerts in this incident, you can zoom in on the critical alert immediately.

With Probable Root Cause, you can cut the time for analysis and shorten the time to resolve!
Here’s the documentation to learn more about the feature. Enjoy!
BASIC | 2 MIN
In this video, you will learn the different incident lists, how to create your own incidents, and how dashboards are created with created incident lists.
Use case walkthrough: Queue and dashboards in APEX AIOps Incident Management ►
This video explains how you can use the queue and dashboards to view incidents in APEX AIOps Incident Management.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
How do you know what’s coming down the pipeline for you to work on in Incident Management? This is the default incident list. This includes all open incidents, regardless of the nature of the issues or assignments.
Your operational procedure may be as simple as just looking at this list and picking an unassigned ticket.

But most likely, you have a queue specific to your team. In this example, we have access to the Application Support team’s incidents view.

So basically this is a list of application-related incidents. Your administrator may have set up a workflow to set the team assignment based on the impacted services, or there may be someone triaging incoming incidents and routing the applicable ones to your team.

And you can make this your default view without affecting other users.

Let’s say we are going to work on this one.

There’s a list view that only shows the incidents you are assigned to.

If you want to filter the list further you can do so here.


You can save the view for yourself without affecting the original view.

If you want to make it a shared view, you can do so here.

When you create a view, you also get a corresponding dashboard.

It presents data in a more visual manner, but you can also jump into a specific incident.

Now you know how to work with your incident queue. Thanks for watching!
BASIC | 2 MIN
In this video, you will learn about the recommendations tab in the Situation Room.
Use case walkthrough: Recommendations tab in Situation Room ►
This video explains how to use the recommendations tab in Situation Room to reference similar incidents, suggest resolving steps, and expedite incident resolution.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
Let’s spend a few minutes learning about the recommendations tab in Situation Room.

Incident Management checks if there are incidents similar to the one at hand. And if there are, it will surface them for you to reference.

In this case, we have one similar incident.

This one is 76% similar to our incident.

And note this icon! This means information on what resolved this incident is available! With this past incident, it looks like the problems were related to a code push. Let’s learn more about it.

Okay, now we have some more context.

Now we can go back to the incident you are working on and see if it’s got that Jenkins alert.
And indeed, here it is. So just like this, the recommendations tab expedites your problem solving.

But how did Incident Management surface that particular incident for us?
How Incident Management identifies similar incidents is configured here.

By default, Incident Management compares these fields and tags to determine similarity. But you can change which fields to use.

And how are resolving steps suggested?
It comes from comments that are tagged as resolving steps.

So as you work on incidents, make sure to always mark the resolving steps. You will be glad you did in the future!

BASIC | 3 MIN
In this video, you will learn more about the top pane in the Situation Room.
Use case walkthrough: Top Pane of Situation Room ►
This video explains how to use the different fields in the top pane of the Situation Room in APEX AIOps Incident Management.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
Let’s take a closer look at each field of the situation room. We’ll focus on the top pane in this video.

This description of the incident is generated by the correlation definition that grouped the alerts.

It is defined here.

So in this example, the location, top three service names, number of sources, and top three source names are all dynamically inserted.

But you can edit it like this, if needed.

This shows the services impacted by this incident.

An incident can be assigned to an individual, and additionally, to one or more groups.

When it’s assigned to a person, the status changes.

How would you know when you have an incident assigned to you? A few ways. Here you can see all the incidents assigned to you.

Or your administrator may have configured an integration to trigger a notification.

The creation time is the time Incident Management grouped these alerts and created an incident. So note that it’s not the time the first event happened.

This shows how long the incident has been open.

If your administrator configured this, you can set a tag or perform tasks using a designated URL.

For example, in our environment you can go here to send this incident to ServiceNow.

Maybe you don’t need to be actively working on this incident, but just want to stay informed. Then click on the watch button.

You can add people other than yourself to watch the incident, too.

Now whenever there’s an announcement about this incident added here, the watchers will receive a notification.

If you want to set priority for your incidents, you can do so here. Then you can sort by priority and tackle the incidents with higher priority.

Now you are familiar with the top section of the situation room. Make sure to check out the other Situation Room deep dive videos!
BASIC | 3 MIN
Take a tour of the Situation Room in APEX AIOps Incident Management!
Use case walkthrough: Tour of the Situation Room ►
This video provides an overview of the Incident Management Situation Room and its features, which include the comments and recommendations tabs.
*Please note Moogsoft is now part of Dell's IT Operations solution called APEX AIOps, and changed its name to APEX AIOps Incident Management. The UI in this video may differ slightly but the content covered is still relevant.
In this video we’ll showcase the power of the Situation Room in Incident Management.
Here’s a critical incident happening…let’s take ownership of it and investigate.

There is a lot going on–several Java Virtual Machines have crashed, and we’re seeing I/O and database problems.

We’ll go to the Situation Room for this incident.
The Situation Room is a virtual collaboration space in Incident Management. It is designed to facilitate collaboration and drive incidents to resolution. Let me show you how it helps our investigation.
It has the same tools and information as the incident details page, but now the entire screen space is dedicated to resolving this one incident. You see the description of the incident, impacted services, and which correlation definition was applied to group the member alerts,

But there are a few things that make the Situation Room special.
First, here is the comments tab. The team can chat as they work through the incident.

Or maybe you just want to monitor the progress of this incident. Then you can add yourself, or a stakeholder as a watcher. You will receive an email when there are any announcements.

Next, the recommendations tab is a great resource. Incident Management checks if there are incidents similar to the one at hand. And if there are, it will surface them for you to reference.
In this case, we have one similar incident.

This one is 76% similar to our incident.

And note this icon! This means information on what resolved this incident is available! With this past incident, it looks like the problems were related to a code push. Let’s learn more about it.

Okay, now we have some more context.

Let’s get back to the incident we were working on. Is there also a Jenkins alert in the current incident? Here’s the timeline that shows you how the incident unfolded…It says code deployment, so this is promising!

Yes, that is the Jenkins alert. It’s likely that this code change is the root cause of the incident. So indeed, that similar incident Incident Management suggested was right!

Let’s share our findings with the team. We’ll talk to the developers and get the deployment rolled back.

All fixed… that was quick! Now you know how the Situation Room supports faster incident resolution. Thanks for watching!

BASIC | 4 MIN
Learn how a user might work in Incident Management. After watching this video, you will be able to identify the typical workflow of a user.
Use case walkthrough: User workflow in APEX AIOps Incident Management ►
This video provides a use case walkthrough of what a typical workflow might look like for an Incident Management user as they resolve incidents.
In this video, we will step through the typical workflow of an APEX AIOps Incident Management user as they work through incidents.
Here comes a slack message, notifying us there’s a critical incident requiring our attention.

So we click through to Incident Management, which takes us to this incident’s Situation Room.
The Situation Room is where you and your team can collaborate on an incident. This is the timeline for this incident. These sliders let you zoom in on particular areas, and the list below filters to match the time frame you choose.

Let’s examine all alerts. These are the alerts that make up this incident. Some of these are alerts from a monitoring system.

And these are alerts generated by Incident Management based on the metrics it is tracking.

We are going to own this incident.

Now we will start our investigation.
These alerts came in within a few seconds of each other. Let’s look at the details of the alert that first came in.

All attributes of this alert are visible now.

And the metric information of the alert is visually presented here.

Incident Management shows you the relevant context and their relationship to each other. This way, it’s much easier to grasp how the whole incident unfolded over time.
According to this, the volume queue length metric exceeded the threshold level and triggered a warning alert.

Then the CPU usage metric on our front end server increased and triggered a warning alert...

...the activity on the backend server fell...

...and we are seeing a backend connection error critical alert.

So, could this be the root cause that had a cascading effect to cause other alerts?

Let’s check out the recommendations tab in the Situation Room for additional insight. Any similar incidents from the past will be surfaced here.

Here's an incident that is 82% similar.

And this icon means it has a resolving step we can review. Great!

This indicates the similar incident was resolved using a runbook tool.

Let’s go to this incident to get more context and confirm we can resolve our incident the same way.

Let’s look at the comments. This incident involved a disk I/O bottleneck that showed up as an increase in Volume Queue Length. Just like our incident.

We can use the same runbook tool to terminate runaway processes and free up resources.

Let’s go back to our incident.
Currently the time window we are seeing is from the moment when the first alert in the incident occurred. We want to see what happens to the metrics when we run the tool. So let’s change the time frame. Now, the metrics are going to be updated in real time.

We’ve run the tool, and with the runaway processes that were overloading I/O killed, the CPU load on the front-end web server is back to normal...

...and activity has resumed on the back end server as well.

Nice! The anomaly has resolved and now the metrics are within the normal range previously learned by the system. Good job!

Now the alerts in our incident are all clear, as well as the incident itself.

The incident status has been changed to resolved. We'll document our solution, and the case is closed!

Just like that, we have resolved our first incident in Incident Management. Now it’s your turn to experience this workflow yourself!
Thanks for watching!







