Microsoft Azure Interview Questions – Part 5: Azure Monitor, Log Analytics, Application Insights & Alerts

Monitoring is one of the most important responsibilities of a senior Azure administrator.

Contents hide
1 Azure Monitor Interview Questions & Answers

Deploying an Azure resource is only the beginning. In production, you need to know:

  • Whether the resource is healthy
  • Whether performance is degrading
  • Why an application is failing
  • Which configuration changed
  • Where an incident started
  • Whether users are affected
  • When to alert administrators
  • How to reduce alert noise
  • How to investigate historical events
  • How to identify reliability and cost improvements

Azure Monitor provides the central observability platform for Azure and hybrid environments. It brings together metrics, logs, traces and events for monitoring and troubleshooting.

This part focuses specifically on monitoring, observability, alerting and operational analysis.

Topics already covered in previous parts, such as VM sizing, VM disk troubleshooting, NSGs, Load Balancer health probes and Application Gateway backend health, are not repeated here except where they are needed to explain monitoring architecture.

Continue the Microsoft Azure Interview Series

← Previous Part: [Part 4: Load Balancer, Application Gateway, Front Door & Traffic Manager] | Complete Series: [Microsoft Azure Interview Questions & Answers – Complete Series]


Azure Monitor Interview Questions & Answers

1. What is Azure Monitor?

Azure Monitor is Microsoft’s unified observability service for Azure and hybrid environments.

It collects and analyzes telemetry such as:

  • Metrics
  • Logs
  • Traces
  • Events

It can then be used for:

  • Monitoring
  • Troubleshooting
  • Alerting
  • Visualization
  • Automation
  • Application performance analysis

Azure Monitor supports both Azure and hybrid resources.


2. What are the major types of telemetry in Azure Monitor?

The important telemetry categories are:

Metrics

Numeric time-series data.

Examples:

CPU percentage
Requests
Latency
Network bytes
Disk operations

Logs

Detailed records of events and activity.

Examples:

Windows events
Application logs
Azure Activity Log
Resource logs

Traces

Used particularly for understanding application request flows and dependencies.

Events

Records of significant occurrences such as resource or configuration changes.

A senior administrator should understand that metrics and logs answer different questions.


3. What is the difference between Azure Monitor Metrics and Logs?

Metrics

Best for:

  • Fast numerical monitoring
  • Dashboards
  • Threshold alerts
  • Time-series analysis

Example:

CPU = 92%

Logs

Best for:

  • Detailed investigation
  • Historical analysis
  • Correlation
  • Searching events
  • KQL queries

Example:

Which application generated the errors?
Which user performed the operation?
What happened immediately before the failure?

A practical approach is:

Metric
  ↓
Detect problem

Logs
  ↓
Investigate problem

Azure Monitor stores logs and metrics in different data platforms.


4. What is a Log Analytics workspace?

A Log Analytics workspace is a centralized data store for Azure Monitor Logs.

It contains tables holding collected log and trace data.

It can receive data from:

  • Azure resources
  • Virtual machines
  • Applications
  • On-premises systems
  • Other supported environments

The data can then be queried using Kusto Query Language (KQL).

 


5. What are the main types of monitoring data in Azure Monitor?

Important telemetry categories include:

Metrics

Numerical measurements collected over time.

Examples:

CPU percentage
Request count
Network throughput
Disk IOPS

Logs

Detailed records of events and activities.

Examples:

Authentication failures
Application exceptions
Azure control-plane operations
Operating system events

Traces

Detailed information about application execution and distributed requests.

Events

Records of activities or changes in state.


6. What is the Azure Activity Log?

Azure Activity Log records Azure control-plane operations.

Examples include:

  • Resource creation
  • Resource deletion
  • Configuration changes
  • Role assignment changes
  • Resource restart operations
  • Administrative operations

A common senior-level question is:

“Who changed this Azure resource?”

The Activity Log is one of the first places to investigate.


7. Is the Activity Log the same as resource logs?

No.

This distinction is important.

Activity Log

Primarily records Azure control-plane operations.

Resource logs

Contain service-specific telemetry generated by an Azure resource.

Therefore:

Activity Log
→ Azure management/control-plane activity

Resource Logs
→ Resource/service-specific telemetry

8. How long is Azure Activity Log data available?

Azure Activity Log data is retained for 90 days by default.

If longer retention or centralized analysis is required, Activity Log events can be exported to supported destinations such as:

  • Log Analytics
  • Storage Account
  • Event Hubs

Retention requirements should be designed according to operational, security and compliance requirements.


9. What is Log Analytics?

Log Analytics is the Azure Monitor experience used to query and analyze log data stored in Log Analytics workspaces.

It uses:

Kusto Query Language (KQL)

Example:

AzureActivity
| where TimeGenerated > ago(24h)
| summarize count() by OperationNameValue

10. What is a Log Analytics workspace?

A Log Analytics workspace is a centralized data store for Azure Monitor logs and traces.

It can receive data from multiple resources.

For example:

VM1 ─────┐
VM2 ─────┤
VM3 ─────┤
App ─────┤
Azure ───┤
          ↓
Log Analytics Workspace
          ↓
        KQL

Organizations may use one or multiple workspaces depending on:

  • Data residency
  • Access boundaries
  • Retention
  • Cost
  • Environment separation
  • Operational ownership

11. What is the difference between a Log Analytics workspace and an Azure Monitor workspace?

They are different Azure resources.

Log Analytics workspace

Primarily used for:

  • Logs
  • Traces
  • KQL-based analysis

Azure Monitor workspace

Used for Azure Monitor metrics scenarios, particularly managed Prometheus metrics.

The two should not be treated as interchangeable.

Log Analytics Workspace
→ Logs / traces
→ KQL

Azure Monitor Workspace
→ Prometheus metrics
→ PromQL

12. What is KQL?

KQL stands for:

Kusto Query Language

It is the query language used to analyze data in Azure Monitor Logs and other Microsoft services.

Example:

AzureActivity
| where TimeGenerated > ago(1h)
| project TimeGenerated, OperationNameValue, ActivityStatusValue, Caller
| order by TimeGenerated desc

KQL is especially important for:

  • Troubleshooting
  • Log analysis
  • Custom alerts
  • Dashboards
  • Security investigation
  • Operational reporting

 

13. What is the basic structure of a KQL query?

A typical query looks like:

TableName
| where Condition
| project Column1, Column2
| summarize count() by Column1
| order by Column1

Example:

AzureActivity
| where TimeGenerated > ago(24h)
| project TimeGenerated, Caller, OperationNameValue, ActivityStatusValue
| order by TimeGenerated desc

14. What does the where operator do?

where filters records.

Example:

AzureActivity
| where ActivityStatusValue == "Failed"

Another example:

AzureActivity
| where TimeGenerated > ago(1h)

15. What does the project operator do?

project selects the columns you want to display.

Example:

AzureActivity
| project TimeGenerated, Caller, OperationNameValue

This makes investigation output easier to read.

16. What does summarize do in KQL?

summarize performs aggregation.

Example:

AzureActivity
| summarize Count=count() by Caller

This can show how many operations were performed by each caller.


17. What does extend do in KQL?

extend creates a calculated column.

Example:

AzureActivity
| extend EventAge = now() - TimeGenerated

It is useful when creating calculated values for analysis.


18. What is the join operator in KQL?

join combines records from two datasets using a common field.

Potential correlation fields include:

  • Resource ID
  • Computer name
  • Request ID
  • Correlation ID

This is particularly useful during advanced investigations where information is distributed across multiple tables.


19. What is the difference between Azure Activity Log and Resource Logs?

This is an important distinction.

Activity Log

Records Azure Resource Manager-level operations.

Examples:

VM created
VM deleted
NSG rule changed
Role assignment changed
Storage Account configuration changed

It is primarily about management-plane activity.

Resource Logs

Provide logs generated by the resource/service itself.

Examples:

Application Gateway access logs
Key Vault logs
Storage logs
Firewall logs

Therefore:

Activity Log
→ What happened to the Azure resource?

Resource Logs
→ What happened inside/at the service?

20. How long is Azure Activity Log retained by default?

Azure Activity Log is retained by Azure for a limited period.

For longer retention or centralized analysis, you can export Activity Log data to destinations such as:

  • Log Analytics workspace
  • Storage Account
  • Event Hub

Retention requirements should therefore be designed rather than relying solely on the default portal history.


21. What is a Diagnostic Setting?

A Diagnostic Setting defines where supported Azure resource logs and metrics should be sent.

Typical destinations include:

  • Log Analytics workspace
  • Storage Account
  • Event Hub
  • Supported partner solutions

Example:

Azure Resource
      |
Diagnostic Setting
      |
+-----+----------+
|     |          |
Logs  Metrics   Events
      |
Log Analytics

Diagnostic settings are an important part of centralized monitoring.


22. What is Azure Monitor Agent?

Azure Monitor Agent (AMA) is the supported Azure Monitor agent for collecting guest operating-system data from Azure and hybrid virtual machines.

It can collect data such as:

  • Windows events
  • Syslog
  • Performance counters
  • Text logs
  • Other supported guest data

The agent uses Data Collection Rules to determine what data to collect and where to send it.


23. What replaced the legacy Log Analytics agent?

The Azure Monitor Agent is the supported agent for guest-OS data collection.

If an organization still has the older Log Analytics agent, it should plan migration according to Microsoft’s current migration guidance.

Do not design new monitoring architectures around the legacy agent.


24. What is a Data Collection Rule (DCR)?

A Data Collection Rule defines:

  • What data to collect
  • How data is processed
  • Where data is sent

For example:

VM
 |
Azure Monitor Agent
 |
DCR
 |
+----------------+
|                |
Windows Events   Performance Counters
 |
Log Analytics Workspace

DCRs provide centralized and consistent monitoring configuration.


25. Why are DCRs important in an enterprise environment?

Imagine an organization has:

1,500 Windows servers
500 Linux servers

Manually configuring every server individually would be difficult to maintain.

With DCRs, you can define standardized collection policies and associate them with appropriate resources.

This provides:

  • Consistency
  • Centralized management
  • Easier changes
  • Better control over data collection
  • Reduced unnecessary ingestion

26. What is a DCR association?

A DCR association connects a resource, such as a VM, to a Data Collection Rule.

Conceptually:

VM
 |
DCR Association
 |
DCR
 |
Data Collection
 |
Destination

A VM can have multiple DCR associations when the monitoring design requires different collection rules.


27. What is data transformation in Azure Monitor?

Data transformation allows collected data to be filtered or modified before it is stored/used.

For example, you may collect a large amount of data but only need certain events.

Conceptually:

Raw Data
   ↓
DCR Transformation
   ↓
Filtered Data
   ↓
Log Analytics

This can help reduce:

  • Unnecessary ingestion
  • Storage
  • Query volume
  • Operational noise

DCRs support data filtering and transformation capabilities.


28. What is VM Insights?

VM Insights is an Azure Monitor capability that simplifies monitoring Azure and hybrid virtual machines.

It helps collect and visualize:

  • Performance data
  • Guest OS information
  • Dependency information where supported/configured
  • VM health-related telemetry

VM Insights can simplify onboarding of Azure Monitor Agent and commonly used performance collection.


29. Does enabling VM Insights mean every possible VM log is collected?

No.

VM Insights provides a simplified monitoring experience and commonly used collection configuration, but monitoring should still be designed around the workload.

You should decide:

  • Which events are needed
  • Which performance counters matter
  • Which logs are valuable
  • Where data should be stored
  • How long data should be retained

Collecting everything without a purpose can increase cost and noise.


30. How would you monitor 1,000 Azure VMs without creating 1,000 separate configurations?

Use scalable monitoring architecture.

A typical approach is:

Azure Policy / Automation
        ↓
Azure Monitor Agent
        ↓
DCR
        ↓
Log Analytics
        ↓
Centralized Alerts

Use resource-based or scope-based alerting where appropriate rather than creating a separate rule for every VM.

Microsoft specifically documents strategies for scaling alert rules across multiple VMs.


31. What is an Azure Monitor alert?

An Azure Monitor alert evaluates monitoring data against defined conditions.

When the condition is met, the alert can trigger an action.

Example:

CPU > threshold
      ↓
Alert Rule
      ↓
Action Group
      ↓
Email / Teams / Automation

Alerts can be based on metrics, logs and other supported signals.


32. What are the main types of Azure Monitor alerts?

Important alert categories include:

  • Metric alerts
  • Log search alerts
  • Activity Log alerts
  • Service Health alerts
  • Resource Health alerts
  • Smart detection/anomaly-related application alerts where applicable

The correct alert type depends on what you are monitoring.


33. What is a metric alert?

A metric alert evaluates a numeric metric.

Example:

CPU > 80%

or:

Available memory < threshold

Metric alerts are generally useful when the required signal is already available as a metric.


34. What is a log search alert?

A log search alert executes a KQL query and evaluates its result.

Example:

AzureActivity
| where ActivityStatusValue == "Failed"
| summarize FailureCount = count()

The alert can trigger when the query result meets the configured condition.

This is useful when the required signal exists in logs rather than as a simple metric.


35. When should you use a metric alert instead of a log alert?

Use a metric alert when:

  • The required metric already exists.
  • You need fast threshold monitoring.
  • You don’t need complex log analysis.

Use a log alert when:

  • The condition depends on log content.
  • You need KQL.
  • Multiple events must be correlated.
  • The required signal is not available as a metric.

Example:

CPU > 90%
→ Metric alert

Five authentication failures from the same IP
→ Log alert

36. What is a dynamic threshold alert?

Dynamic thresholds use historical behavior and machine-learning techniques to determine expected metric or query behavior.

Instead of manually specifying:

CPU > 80%

the system can learn normal patterns such as:

Monday 9 AM → High
Monday 2 AM → Low
Weekend → Low

and identify deviations.

Dynamic thresholds are useful when normal values vary over time.


37. When can static thresholds be better than dynamic thresholds?

Static thresholds are often better when the limit is based on a hard operational requirement.

Example:

Disk free space < 10%

or:

Certificate expiration < 14 days

If the business requirement is fixed, a static threshold can be easier to understand and operate.


38. What is an Action Group?

An Action Group defines what should happen when an Azure Monitor alert fires.

Actions can include:

  • Email
  • SMS
  • Push notification
  • Voice
  • Webhook
  • Azure Function
  • Logic App
  • Other supported automation mechanisms

Action Groups are reusable and can be associated with multiple alerts.


39. Why should Action Groups be separated from alert rules?

Suppose an organization has:

100 alert rules

and all need to notify the same operations team.

Instead of configuring recipients separately in every rule:

100 Alerts
     ↓
One reusable Action Group
     ↓
Operations Team

This simplifies administration.

It also allows notification destinations to be changed centrally.


40. Can one alert use multiple Action Groups?

Yes.

Azure Monitor allows multiple action groups to be associated with an alert rule.

The actions are executed concurrently rather than in a guaranteed sequence.


41. What is alert fatigue?

Alert fatigue occurs when administrators receive too many alerts.

Example:

10,000 alerts/month
       ↓
Most are low-value
       ↓
Administrators ignore alerts
       ↓
Critical alert gets missed

A good monitoring design prioritizes:

  • Actionable alerts
  • Appropriate thresholds
  • Deduplication
  • Suppression where appropriate
  • Meaningful severity
  • Correct recipients

42. How would you reduce alert noise?

Use:

  1. Meaningful thresholds.
  2. Dynamic thresholds where appropriate.
  3. Aggregation.
  4. Scope-based alerts.
  5. Alert processing rules where appropriate.
  6. Maintenance suppression.
  7. Correct severity.
  8. Appropriate evaluation frequency.
  9. KQL filtering.
  10. Different notification paths for different severities.

The objective is:

Fewer alerts
+
Higher signal
=
Better operations

43. What is an alert processing rule?

Alert processing rules allow you to modify how alerts are processed after they are generated.

Depending on the supported scenario, they can be used to:

  • Add/remove action groups
  • Suppress notifications
  • Apply rules during maintenance windows

This can be useful when the alert itself remains valid but notification behavior needs temporary modification.


44. What is the difference between disabling an alert and suppressing its notifications?

Disable alert

Stops the alert rule from evaluating/generating alerts.

Suppress notification

The alert can still be generated, but notification behavior is changed.

This distinction is useful during planned maintenance.

For example:

Maintenance Window
        ↓
Keep alert rule active
        ↓
Suppress notifications

This preserves monitoring while avoiding unnecessary notifications.


45. What is Azure Service Health?

Azure Service Health provides information about Azure service incidents, planned maintenance and health advisories that can affect your resources or services.

It includes experiences such as:

  • Service issues
  • Planned maintenance
  • Health advisories

It is particularly useful during Azure platform incidents.


46. What is Azure Resource Health?

Resource Health provides information about the health of an individual Azure resource.

Example:

Azure VM
   ↓
Resource Health
   ↓
Healthy / Degraded / Unavailable

This is different from Service Health.


47. What is the difference between Service Health and Resource Health?

Service Health

Answers:

Is an Azure service/platform issue affecting me or my environment?

Resource Health

Answers:

What is the health state of this specific resource?

Example:

Service Health
→ Azure Storage regional incident

Resource Health
→ This particular VM is unavailable

Both should be checked during major incidents.


48. What is the Azure Activity Log useful for during an incident?

Suppose a production application suddenly stops working.

You discover:

NSG rule changed 10 minutes ago

Activity Log can help identify the management-plane operation.

You can investigate:

  • Who performed it
  • What operation occurred
  • Which resource was affected
  • When it happened
  • Whether the operation succeeded or failed

This makes Activity Log extremely useful for change-related troubleshooting.


49. How would you investigate an unexplained configuration change?

Use:

Activity Log
     ↓
Time of change
     ↓
Operation
     ↓
Caller
     ↓
Resource
     ↓
Related deployment/change

Then correlate with:

  • Azure Resource Graph
  • Change history where available
  • Deployment records
  • Source-control/IaC changes
  • Administrative records

50. What is Azure Monitor Workbook?

A Workbook is an interactive reporting and visualization tool in Azure Monitor.

It can combine:

  • Text
  • Metrics
  • KQL queries
  • Charts
  • Parameters
  • Tables

Example:

Operations Workbook
│
├── VM health
├── CPU
├── Memory
├── Storage
├── Application errors
└── Network metrics

This is useful for creating operational dashboards.


51. What is the difference between a Dashboard and a Workbook?

Azure Dashboard

Primarily a customizable portal dashboard for displaying resource information and visualizations.

Workbook

Designed for interactive monitoring and reporting using:

  • Queries
  • Metrics
  • Parameters
  • Visualizations
  • Text

Workbooks are particularly useful for building operational investigation views.


52. What is Application Insights?

Application Insights is an application performance monitoring capability within Azure Monitor.

It helps monitor applications by collecting telemetry such as:

  • Requests
  • Dependencies
  • Exceptions
  • Traces
  • Performance information
  • Availability

Current Azure Monitor documentation describes Application Insights as an OpenTelemetry-based application monitoring capability.


53. What is the difference between Azure Monitor and Application Insights?

Think of it as:

Azure Monitor
      |
      +--- Infrastructure monitoring
      |
      +--- Metrics
      |
      +--- Logs
      |
      +--- Alerts
      |
      +--- Application Insights
                |
                +--- Application telemetry

Application Insights focuses on application performance and behavior, while Azure Monitor provides the broader observability platform.


54. What is distributed tracing?

Distributed tracing follows a request as it moves across multiple application components.

Example:

User
 ↓
Web App
 ↓
API
 ↓
Service
 ↓
Database

If the total request takes 4 seconds, distributed tracing can help determine which dependency consumed most of the time.

This is especially valuable for microservices.


55. What is an Application Insights dependency?

A dependency is an external component that an application calls.

Examples:

  • Database
  • HTTP API
  • Storage
  • Queue
  • Other services

Example:

Web Application
      ↓
SQL Database
      ↓
Dependency

Dependency telemetry helps identify slow or failing downstream components.


56. An application is slow. CPU is normal. What would you investigate?

Do not assume the application itself is healthy simply because CPU is low.

Use Application Insights to investigate:

  • Request duration
  • Dependency duration
  • Failed requests
  • Exceptions
  • External API latency
  • Database latency
  • Trace information

Example:

Request = 4 seconds

Application processing = 200 ms
Database = 3.5 seconds

The application server CPU can be normal while the application remains slow.


57. How can Application Insights help identify a database bottleneck?

Analyze dependency telemetry.

Example:

HTTP Request
   |
   +--- Application processing: 100 ms
   |
   +--- SQL dependency: 2.8 sec

This points the investigation toward the database rather than immediately scaling the application server.


58. What is an Application Insights availability test?

Availability testing checks whether an application endpoint is reachable and responding as expected.

It can help detect:

  • Application outage
  • Endpoint failure
  • Excessive response time
  • Regional availability problems

Availability tests can generate alerts when the endpoint becomes unavailable.


59. Why are availability tests useful if you already have server monitoring?

Server monitoring answers:

Is my server healthy?

Availability testing answers:

Can a user actually reach the application and receive an expected response?

These are different.

Example:

VM CPU → 20%
VM Memory → 40%
VM → Healthy

Website → HTTP 500

Infrastructure monitoring alone may not detect the application-level outage.


60. What is synthetic monitoring?

Synthetic monitoring uses automated requests to simulate user interactions or endpoint access.

For example:

Every 5 minutes:

Open website
     ↓
Check response
     ↓
Check response time
     ↓
Alert if failure

This provides an external perspective of application availability.


61. How would you troubleshoot an application that users report as slow?

Use a top-down approach:

User
 ↓
Application availability
 ↓
Request duration
 ↓
Application processing
 ↓
Dependencies
 ↓
Database/API/Storage
 ↓
Infrastructure

Use:

  • Application Insights
  • Azure Monitor Metrics
  • Log Analytics
  • Application logs

Avoid starting with VM resizing unless telemetry indicates a compute bottleneck.


62. What is a correlation ID and why is it useful?

A correlation ID is an identifier used to associate related operations across components.

Example:

Request ID:
abc-123

The same identifier can be included in:

Web application
 ↓
API
 ↓
Database/logging

During troubleshooting, it helps correlate events belonging to the same user request.


63. How would you investigate an HTTP 500 application error?

Start with Application Insights.

Check:

  1. Failed requests.
  2. Exceptions.
  3. Request details.
  4. Dependencies.
  5. Traces.
  6. Deployment/change history.
  7. Application logs.

Then correlate the failure timestamp with infrastructure telemetry.


64. What is Azure Monitor Logs cost optimization?

Log ingestion and retention can generate significant costs in large environments.

Optimization strategies include:

  • Collect only required data.
  • Filter unnecessary events.
  • Use DCR transformations.
  • Select appropriate table plans.
  • Review retention.
  • Avoid duplicate collection.
  • Monitor workspace usage.

DCR-based filtering and transformations can help control unnecessary ingestion.


65. Why is collecting every Windows Event Log channel not always a good idea?

Because large environments can generate huge volumes of data.

For example:

5,000 servers
×
Large event volume
=
Very high ingestion

This can increase:

  • Cost
  • Query complexity
  • Noise
  • Storage requirements

Instead, identify which events are actually required for:

  • Operations
  • Security
  • Compliance
  • Troubleshooting

66. How would you design Log Analytics workspaces for a large organization?

There is no universal rule such as:

“Always use one workspace.”

Evaluate:

  • Geographic requirements
  • Data residency
  • Security boundaries
  • Operational teams
  • Retention
  • Cost
  • Query requirements
  • Microsoft Sentinel architecture
  • Cross-resource monitoring

Some organizations use centralized workspaces, while others use multiple workspaces for isolation and governance.


67. What is the difference between Log Analytics workspace and Azure Monitor workspace?

This is a current and important Azure terminology question.

Log Analytics workspace

Used for Azure Monitor Logs:

  • Logs
  • Traces
  • KQL

Azure Monitor workspace

Currently used for Prometheus metrics collected by Azure Monitor.

It is a different resource type with a different data platform.

Do not treat the two as interchangeable merely because both names contain “workspace.”


68. What is Prometheus?

Prometheus is a monitoring and metrics system commonly used for infrastructure and cloud-native workloads.

It stores numeric time-series metrics.

Example:

http_requests_total
cpu_usage
memory_usage

Azure Monitor supports managed Prometheus scenarios.

Prometheus metrics are currently stored in Azure Monitor workspaces and queried using PromQL.


69. What is PromQL?

PromQL stands for:

Prometheus Query Language

It is used to query Prometheus metrics.

This differs from KQL:

KQL
→ Azure Monitor Logs

PromQL
→ Prometheus metrics

This distinction is increasingly important in modern Azure monitoring environments.


70. What is Azure Advisor?

Azure Advisor is a recommendation service that analyzes Azure resource configuration and usage telemetry and provides recommendations across:

  • Reliability
  • Security
  • Performance
  • Cost
  • Operational Excellence

It is intended to help optimize Azure deployments rather than act as a real-time monitoring system.


71. Is Azure Advisor the same as Azure Monitor?

No.

Azure Monitor

Answers:

What is happening now and what happened?

Azure Advisor

Answers:

What improvements does Azure recommend for my environment?

Example:

Azure Monitor
→ CPU has been high

Azure Advisor
→ Consider changing configuration/SKU based on observed usage

They complement each other.


72. What are the five Azure Advisor categories?

Current Advisor categories are:

  1. Reliability
  2. Security
  3. Performance
  4. Cost
  5. Operational Excellence

These categories cover different optimization areas.


73. Should you blindly implement every Azure Advisor recommendation?

No.

Advisor recommendations are recommendations, not automatic architectural decisions.

For each recommendation, evaluate:

  • Business requirements
  • Application architecture
  • Risk
  • Cost
  • Availability requirements
  • Security requirements
  • Maintenance impact

A recommendation that is appropriate for one workload may not be appropriate for another.


74. What is Azure Advisor useful for during a cost optimization exercise?

Advisor can identify opportunities such as:

  • Idle resources
  • Underutilized resources
  • Appropriate sizing opportunities
  • Storage-related cost improvements
  • Other cost optimization opportunities

Cost recommendations are available through the Cost category.


75. What is the difference between Azure Advisor and Azure Cost Management?

Azure Advisor

Provides recommendations.

Azure Cost Management

Provides detailed cost analysis, budgets, cost allocation and financial management capabilities.

Example:

Advisor
→ "This resource may be underutilized."

Cost Management
→ "This subscription spent ₹X / $X and this resource group consumed Y%."

They solve different operational problems.


76. How would you investigate unexpected Azure cost growth?

Use a combination of:

Cost Management
+
Azure Advisor
+
Azure Monitor
+
Resource inventory

Check:

  • Which service increased?
  • Which resource increased?
  • Was there a deployment?
  • Did data transfer increase?
  • Did log ingestion increase?
  • Did VM capacity increase?
  • Did storage grow?
  • Did a resource stop being deallocated?

Do not assume that the most expensive resource is automatically the cause of the increase.


77. How can excessive Azure Monitor logging increase cost?

Consider:

More data collected
      ↓
More ingestion
      ↓
More storage
      ↓
More query/retention requirements
      ↓
Higher cost

Therefore, monitoring itself must be designed economically.


78. How would you monitor a production Azure application end-to-end?

A good architecture could be:

Users
  |
Application
  |
Application Insights
  |
+-------------------------+
|                         |
Requests             Dependencies
|                         |
Exceptions             Database
|                         |
Traces                  APIs
|
Azure Monitor
|
+----------------------------+
|                            |
Metrics                   Logs
|                            |
Alerts                   Log Analytics
|                            |
Action Groups             KQL

Then use Azure Advisor for periodic optimization recommendations.


79. How would you design monitoring for an on-premises + Azure hybrid environment?

Use a common observability architecture.

Example:

Azure VMs
     \
      \
On-Prem Servers → Azure Monitor
      /
Azure Resources

Use:

  • Azure Monitor Agent
  • DCRs
  • Log Analytics
  • Azure Arc where appropriate
  • Centralized alerts
  • Application Insights

Azure Monitor supports hybrid monitoring scenarios.


80. A monitoring agent is installed but no VM logs appear. What would you check?

Use this sequence:

VM
 ↓
Azure Monitor Agent
 ↓
DCR Association
 ↓
DCR Data Source
 ↓
DCR Destination
 ↓
Network
 ↓
Log Analytics

Check:

  1. Agent installation.
  2. Agent health.
  3. DCR association.
  4. Data source configuration.
  5. Destination.
  6. Workspace permissions/configuration.
  7. Network connectivity.
  8. Whether the expected data is actually being generated.

Do not reinstall the agent immediately.


81. A DCR exists but no data arrives. What would you check?

Check:

DCR
 |
+-- Data source
 |
+-- Destination
 |
+-- Transformation
 |
+-- Association

Specifically verify:

  • Correct VM/resource association.
  • Correct event/performance source.
  • Correct destination workspace.
  • No transformation filtering the required records.
  • Agent is running.
  • Expected events are being generated.

82. Logs appear in Log Analytics but an alert does not fire. What would you investigate?

Check:

  1. KQL query.
  2. Time range.
  3. Alert evaluation frequency.
  4. Alert condition.
  5. Threshold.
  6. Query result.
  7. Alert rule scope.
  8. Action Group.
  9. Alert processing rules.
  10. Notification channel.

A useful approach is:

KQL query works manually?
        |
       YES
        ↓
Alert condition correct?
        ↓
Action Group correct?
        ↓
Notification path working?

83. An alert fires repeatedly every few minutes. How would you reduce the noise?

Check:

  • Evaluation frequency
  • Threshold
  • Aggregation
  • Dynamic threshold
  • Alert dimensions
  • Suppression/processing rules
  • Whether the alert condition clears properly
  • Whether one alert rule is creating many independent alert instances

The goal is to alert on actionable conditions rather than every individual telemetry fluctuation.


84. CPU frequently crosses 80% for only one minute. Should you alert?

Not necessarily.

Ask:

  • Is the spike expected?
  • Is it correlated with application impact?
  • Is CPU sustained?
  • Is the application actually degraded?

A better alert might be:

CPU > 80%
for 10 minutes

rather than:

CPU > 80%
for 1 minute

The correct threshold and duration depend on the workload.


85. How would you monitor a critical website from the user’s perspective?

Use:

  • Application Insights availability tests
  • Application Insights request telemetry
  • Dependency telemetry
  • Azure Monitor alerts

Monitor:

Availability
Response time
HTTP failures
Dependencies
Exceptions

This detects problems that infrastructure-only monitoring may miss.


86. How would you investigate a website that is available but slow?

Compare:

Availability
      ↓
Request duration
      ↓
Dependency duration
      ↓
Exceptions
      ↓
Infrastructure metrics

For example:

Availability = 100%

Request latency = 5 sec

Database dependency = 4.2 sec

The issue is likely in a dependency rather than basic website availability.


87. How would you investigate an outage that began immediately after a deployment?

Correlate:

Deployment time
      ↓
Application errors
      ↓
Application Insights
      ↓
Activity Log
      ↓
Configuration changes
      ↓
Dependency failures

Determine whether:

  • Code changed
  • Configuration changed
  • Identity permissions changed
  • Network configuration changed
  • Secret/certificate changed
  • Dependency version changed

This is much more reliable than simply restarting the application.


88. How would you investigate a production incident where users report intermittent failures?

Start by determining:

All users?
Some users?
One region?
One backend?
One API?
One dependency?

Then correlate:

  • Request IDs
  • Application Insights
  • Backend health
  • Logs
  • Metrics
  • Deployment history
  • Activity Log

Intermittent problems often require correlation across multiple telemetry sources.


89. What is the role of Azure Monitor in incident response?

Azure Monitor provides:

Detection
   ↓
Investigation
   ↓
Alerting
   ↓
Automation
   ↓
Evidence for root cause analysis

It should be part of an operational process rather than simply a dashboard.


90. How can Azure Monitor trigger automated remediation?

An alert can invoke an Action Group that triggers supported automation such as:

  • Azure Function
  • Logic App
  • Webhook
  • Other supported actions

Example:

Service failure
     ↓
Azure Monitor Alert
     ↓
Action Group
     ↓
Logic App
     ↓
Remediation workflow

Action Groups support automated actions in addition to human notifications.


91. Give an example of safe automated remediation.

Example:

Application service stopped
        ↓
Alert
        ↓
Action Group
        ↓
Automation
        ↓
Restart service

However, automation should include:

  • Scope controls
  • Failure handling
  • Logging
  • Idempotency
  • Approval where appropriate

Blind automation can make incidents worse.


92. What would you include in an enterprise Azure monitoring strategy?

I would divide it into several layers:

Platform

  • Service Health
  • Resource Health
  • Activity Log

Infrastructure

  • VM metrics
  • Guest OS telemetry
  • Storage
  • Network

Application

  • Requests
  • Dependencies
  • Exceptions
  • Traces
  • Availability

Security

  • Relevant security logs
  • Microsoft Defender for Cloud
  • Microsoft Sentinel where applicable

Alerting

  • Action Groups
  • Severity
  • Escalation
  • Suppression

Operations

  • Workbooks
  • Dashboards
  • KQL
  • Runbooks/automation

Optimization

  • Azure Advisor
  • Cost Management

93. What is your complete approach to designing production monitoring?

A strong senior-level answer is:

First, identify the business-critical services
and their SLO/SLA requirements.

Then define what must be measured:

Availability
Latency
Errors
Capacity
Dependencies
Security events
Configuration changes

94. Why would you filter logs before ingestion?

Suppose thousands of resources generate verbose logs that nobody uses.

Ingesting everything can increase:

  • Data volume
  • Storage
  • Query workload
  • Monitoring cost

Filtering genuinely unnecessary telemetry can therefore improve both operational efficiency and cost control.

However, filtering should be based on documented requirements rather than simply deleting data because it appears unimportant.


95. What are Diagnostic Settings?

Diagnostic settings configure supported Azure resources to send platform logs and metrics to destinations such as:

  • Log Analytics workspace
  • Storage Account
  • Event Hubs
  • Supported partner solutions

Conceptually:

Azure Resource
      ↓
Diagnostic Setting
      ↓
Destination

Diagnostic settings are therefore an important part of building centralized monitoring.


96. Are resource logs automatically available for every Azure resource?

No.

A common mistake is assuming that because Azure Monitor exists, every detailed resource log is automatically stored.

For many Azure services, you must configure Diagnostic Settings to send resource logs to a destination.

Therefore, if an administrator cannot find historical resource logs, check:

  1. Whether the resource supports the required log category
  2. Whether Diagnostic Settings were configured
  3. Whether the correct destination was selected
  4. Whether the expected time range is covered
  5. Whether the relevant logs were actually generated

97. How would you troubleshoot a resource where expected logs are missing?

Use a structured approach:

Step 1

Identify the resource.

Step 2

Check Diagnostic Settings.

Step 3

Verify enabled categories.

Step 4

Verify destination.

Step 5

Check the destination workspace.

Step 6

Confirm the expected table exists.

Step 7

Run a time-bounded KQL query.

Step 8

Check whether the resource actually generated the event.

Do not immediately assume that Log Analytics itself is broken.


98. What is Application Insights?

Application Insights is an Azure Monitor application performance monitoring capability.

It helps monitor application behavior such as:

  • Requests
  • Dependencies
  • Exceptions
  • Traces
  • Availability
  • Performance
  • Distributed transactions

It is designed to answer questions such as:

“The application is slow. Which dependency or operation is causing the delay?”


99. What is workspace-based Application Insights?

Workspace-based Application Insights stores application telemetry in a Log Analytics workspace.

This allows application telemetry to be queried using KQL alongside other Azure Monitor logs.

For example:

Application
    ↓
Application Insights
    ↓
Log Analytics Workspace
    ↓
KQL

This is particularly useful for centralized observability.


100. What types of application telemetry can Application Insights collect?

Common telemetry includes:

  • Requests
  • Dependencies
  • Exceptions
  • Traces
  • Page views
  • Custom events
  • Application metrics

For workspace-based Application Insights, commonly used tables include:

AppRequests
AppDependencies
AppExceptions
AppTraces
AppPageViews

The exact telemetry available depends on the application and instrumentation.


101. What is distributed tracing?

Distributed tracing follows a request as it travels through multiple application components.

For example:

User
 ↓
Web Application
 ↓
API
 ↓
Database
 ↓
External Service

A distributed trace can help identify which component contributed to latency or failure.

This is particularly valuable in microservice environments.


102. What is Application Map in Application Insights?

Application Map provides a visual representation of application components and their dependencies.

For example:

Web App
   |
   +---- API
   |
   +---- SQL Database
   |
   +---- External API

It can help identify:

  • Failed components
  • Slow dependencies
  • Request relationships
  • Application architecture

It is useful as an investigation starting point, but detailed analysis should still use telemetry and logs.


103. What are Application Insights availability tests?

Availability tests allow you to test whether an application endpoint is reachable and responding as expected from supported test locations.

They help answer:

“Can users reach the application?”

They can detect availability problems independently of normal user traffic.


104. How would you investigate a sudden increase in application exceptions?

Start with the time range when exceptions increased.

Then investigate:

Exception count
      ↓
Exception type
      ↓
Affected operation
      ↓
Affected dependency
      ↓
Deployment/configuration change

Example KQL:

AppExceptions
| where TimeGenerated > ago(2h)
| summarize Count=count() by ExceptionType
| order by Count desc

Then investigate the most significant exception types in context.


105. How would you identify slow application requests using KQL?

For example:

AppRequests
| where TimeGenerated > ago(1h)
| where DurationMs > 2000
| project TimeGenerated, Name, DurationMs, ResultCode
| order by DurationMs desc

This identifies requests taking more than two seconds.

The threshold should be adjusted according to the application’s actual SLA/SLO.


106. How would you identify failed application requests?

Example:

AppRequests
| where TimeGenerated > ago(1h)
| where Success == false
| summarize Failures=count() by Name
| order by Failures desc

This helps identify which operations are producing the most failures.


107. How would you investigate dependencies causing application failures?

Example:

AppDependencies
| where TimeGenerated > ago(1h)
| where Success == false
| summarize Failures=count() by Target, DependencyType
| order by Failures desc

You can then correlate the dependency failures with:

  • Application exceptions
  • Request failures
  • Deployment changes
  • Infrastructure events

108. What is an Azure Monitor metric alert?

A metric alert triggers when a metric meets a configured condition.

Example:

Metric:
CPU Percentage

Condition:
Greater than 90%

Evaluation:
5 minutes

Action:
Send notification

Metric alerts are useful when you need fast notification based on numerical telemetry.


109. What is a log search alert?

A log search alert evaluates a KQL query and triggers when the query result meets the configured condition.

Example:

AzureActivity
| where ActivityStatusValue == "Failed"

You could configure an alert based on the number of matching events.

Log-based alerts are useful when the condition cannot be represented adequately by a simple metric.


110. What is an Activity Log alert?

An Activity Log alert triggers when a specified Azure Activity Log event occurs.

For example:

Administrative operation
+
Specific resource
+
Specific operation

A security team could use an Activity Log alert for sensitive administrative changes.


111. What is the difference between metric alerts and log alerts?

Metric alert

Works against numerical metric data.

Example:

CPU > 90%

Log alert

Evaluates records using KQL.

Example:

Count of failed administrative operations > 5

A simple way to remember:

Metric alert → numerical condition

Log alert → query-based condition

112. What are dynamic thresholds in Azure Monitor alerts?

Dynamic thresholds allow Azure Monitor to establish expected behavior based on historical patterns rather than requiring a manually fixed threshold in every situation.

For example, instead of:

CPU > 80%

the monitoring system can learn expected patterns and identify unusual deviations.

Dynamic thresholds can be useful when:

  • Workload changes throughout the day
  • Traffic is seasonal
  • Different resources have different normal behavior
  • Static thresholds produce excessive alerts

They should still be validated against the actual business workload.


113. What is an Action Group?

An Action Group defines what happens when an alert is triggered.

Possible actions can include:

  • Email
  • SMS
  • Push notifications
  • Voice notifications
  • Webhooks
  • Azure Functions
  • Logic Apps
  • Automation-related actions

Example:

Alert
 ↓
Action Group
 ↓
Email IT Team
+
Send webhook
+
Trigger automation

Action Groups allow the notification mechanism to be reused by multiple alert rules.


114. Why should Action Groups be standardized?

Suppose an organization has:

500 alert rules

If every alert has individually configured recipients, changing an operations team’s contact information becomes difficult.

Instead:

Critical-Production
        ↓
Action Group
        ↓
Production Team

Multiple alert rules can reference the same Action Group.

This simplifies administration and change management.


115. What are Alert Processing Rules?

Alert Processing Rules allow organizations to modify how alerts are processed after they are generated.

They can be used for scenarios such as:

  • Adding action groups
  • Suppressing notifications
  • Scheduled maintenance windows
  • Resource-specific alert handling

For example:

01:00–03:00
Planned maintenance

↓
Suppress selected notifications

This can reduce unnecessary alert noise during planned activities.


116. What is alert fatigue?

Alert fatigue occurs when administrators receive too many alerts, including alerts that do not require action.

For example:

1000 alerts/day
↓
950 are ignored
↓
50 are investigated

This is dangerous because important alerts can be overlooked.

A good monitoring design should prioritize:

  • Actionable alerts
  • Correct thresholds
  • Alert severity
  • Deduplication
  • Appropriate notification channels
  • Maintenance suppression
  • Clear ownership

117. How would you reduce excessive Azure Monitor alerts?

I would use this process:

1. Identify noisy alerts

Determine which alerts fire most frequently.

2. Review thresholds

Check whether thresholds reflect actual operational requirements.

3. Remove duplicate alerts

Avoid multiple alerts for the same condition.

4. Use dynamic thresholds where appropriate

This can help for variable workloads.

5. Use alert processing rules

Suppress notifications during approved maintenance windows.

6. Route alerts to the correct team

Not every alert should go to every administrator.

7. Review alerts periodically

Monitoring configuration should evolve with the environment.


118. What is Azure Service Health?

Azure Service Health provides information about Azure service-related events that may affect your resources.

Important categories include:

  • Service issues
  • Planned maintenance
  • Health advisories

It can help determine whether an incident is related to an Azure platform event rather than an organization’s own configuration.


119. What is the difference between Service Health, Resource Health and Activity Log?

This is a common senior-level interview question.

Service Health

Answers:

Is there an Azure platform event that may affect my services?

Resource Health

Answers:

What is the health state of this particular Azure resource?

Activity Log

Answers:

What Azure control-plane operations occurred?

Therefore:

Service Health
→ Azure platform/service events

Resource Health
→ Individual resource health

Activity Log
→ Management/control-plane activity

120. What is Azure Monitor Workbooks?

Azure Monitor Workbooks provide interactive reporting and visualization.

They can combine:

  • Metrics
  • Logs
  • KQL queries
  • Parameters
  • Charts
  • Tables
  • Text
  • Other Azure monitoring information

For example:

Production Monitoring Workbook

VM Health
Application Errors
Request Latency
Failed Operations
Security Events
Resource Health

A workbook can therefore provide an operational view without requiring administrators to run individual queries repeatedly.


121. What is the difference between Azure dashboards and Workbooks?

Dashboards

Primarily provide a customizable collection of tiles.

Workbooks

Provide richer interactive reporting capabilities, including:

  • Queries
  • Parameters
  • Filters
  • Multiple visualizations
  • Narrative text
  • Interactive analysis

A senior administrator might use:

Dashboard
→ High-level operational status

Workbook
→ Detailed investigation and analysis

122. What is Azure Advisor?

Azure Advisor analyzes Azure resource configuration and usage and provides recommendations.

Advisor recommendations can cover areas such as:

  • Reliability
  • Security
  • Performance
  • Cost
  • Operational Excellence

Examples may include recommendations related to:

  • Underutilized resources
  • Security configuration
  • Resiliency
  • Performance optimization
  • Cost optimization

Advisor should be treated as a recommendation engine, not as an automatic authority. Administrators should validate recommendations against application requirements before implementing them.


123. How can Azure Advisor help with cost optimization?

Advisor can identify opportunities such as underutilized or oversized resources.

For example:

VM allocated:
8 vCPU

Observed workload:
Low utilization

Advisor may recommend reviewing the resource configuration.

Before changing it, verify:

  • Business requirements
  • Peak utilization
  • Performance requirements
  • Availability requirements
  • Licensing implications
  • Application dependencies

124. What is Azure Cost Management?

Azure Cost Management provides tools to analyze and control Azure spending.

Important capabilities include:

  • Cost Analysis
  • Budgets
  • Cost alerts
  • Cost allocation
  • Forecasting
  • Cost optimization analysis

A senior administrator should understand that monitoring is not limited to availability and performance.

Cost is also an operational signal.


125. How would you investigate an unexpected increase in Azure cost?

I would follow this process:

Cost increase detected
        ↓
Identify subscription
        ↓
Identify resource group
        ↓
Identify resource/service
        ↓
Compare with previous period
        ↓
Identify usage increase
        ↓
Check configuration changes
        ↓
Check deployments
        ↓
Check data transfer / ingestion
        ↓
Take corrective action

For monitoring specifically, excessive Log Analytics ingestion can itself become a significant cost factor.


126. How can monitoring itself become expensive?

Monitoring can generate significant costs when an organization collects unnecessary amounts of telemetry.

Potential causes include:

  • Excessive log ingestion
  • Verbose application logging
  • Unnecessary diagnostic categories
  • Long retention periods
  • High-volume telemetry
  • Repeated data collection
  • Poorly designed monitoring

A good monitoring architecture balances:

Observability
+
Security
+
Troubleshooting
+
Cost

The goal is not to collect everything indiscriminately.


127. How would you design monitoring for 500 Azure VMs?

I would avoid configuring every VM independently.

A scalable design could look like:

                Azure Monitor
                     |
        +------------+-------------+
        |                          |
     Metrics                    Logs
        |                          |
     Alerts                Log Analytics
                                   |
                                  KQL
                                   |
                              Workbooks
                                   |
                            Incident Response

For VM log collection:

VMs
 ↓
Azure Monitor Agent
 ↓
DCRs
 ↓
Log Analytics

Then I would implement:

  • Standardized monitoring policies
  • Appropriate DCRs
  • Centralized workbooks
  • Action Groups
  • Alert severity standards
  • Service Health alerts
  • Resource Health monitoring
  • Cost monitoring
  • Azure Advisor review

128. A production application suddenly becomes slow. How would you investigate it using Azure Monitor?

I would not immediately restart resources.

I would establish:

Step 1 — Define the incident

Determine:

  • Start time
  • End time
  • Affected users
  • Affected application
  • Business impact

Step 2 — Check application telemetry

Review:

  • Request duration
  • Request failures
  • Exceptions
  • Dependencies

Step 3 — Check recent changes

Look for:

  • Application deployment
  • Configuration change
  • Infrastructure change
  • Identity change

Step 4 — Correlate telemetry

Use:

Requests
+
Dependencies
+
Exceptions
+
Activity Log

Step 5 — Determine the failing component

For example:

Requests slow
      ↓
Dependency latency increased
      ↓
Database response increased

The objective is to identify the actual cause rather than simply treating the symptom.


129. An alert says a VM is unavailable, but the application team says users are still working. What do you do?

Do not immediately close the alert.

First determine:

  1. Which alert rule fired?
  2. Which metric or condition triggered it?
  3. What resource was evaluated?
  4. What was the evaluation period?
  5. Was the condition transient?
  6. Is the application actually affected?
  7. Is the alert threshold appropriate?

Then compare:

Alert telemetry
+
Resource health
+
Application telemetry
+
User impact

An alert is an indication requiring investigation; it is not automatically proof of business impact.


130. A critical alert did not fire during an incident. How would you troubleshoot the monitoring configuration?

I would check:

1. Alert rule

Was it enabled?

2. Scope

Was the affected resource included?

3. Condition

Did the telemetry actually meet the condition?

4. Evaluation period

Was the condition sustained long enough?

5. Data availability

Was telemetry being collected?

6. Diagnostic configuration

If the alert depended on logs, were the relevant logs reaching the workspace?

7. Action Group

Was the notification/action configuration correct?

8. Alert processing rules

Was notification suppressed?

This separates:

Alert condition problem

from:

Notification/action problem

131. What would you check if Azure Monitor Agent is installed but expected VM logs are missing?

I would check:

VM
 ↓
AMA installed/running?
 ↓
DCR associated?
 ↓
DCR collecting required data?
 ↓
Correct destination?
 ↓
Network/connectivity requirements?
 ↓
Expected table receiving data?

I would also check whether the requested data source is actually configured in the DCR.

The presence of the agent alone does not prove that the required telemetry is being collected.


132. Describe your approach to a major Azure production monitoring incident.

My approach would be:

Phase 1 — Establish impact

Identify:

  • What is failing?
  • Who is affected?
  • When did it begin?
  • Is it still occurring?

Phase 2 — Establish timeline

Correlate:

Application telemetry
        +
Azure metrics
        +
Activity Log
        +
Resource Health
        +
Service Health
        +
Recent changes

Phase 3 — Identify the failing layer

Determine whether the problem is primarily:

Application
Identity
Platform
Resource
Configuration
Dependency
Monitoring

Phase 4 — Contain

Apply the minimum safe action required to restore service.

Phase 5 — Validate

Confirm recovery through:

  • Metrics
  • Logs
  • Application telemetry
  • Health status
  • User/business validation

Phase 6 — Prevent recurrence

After recovery:

  • Tune alerts
  • Improve dashboards/workbooks
  • Improve logging
  • Review DCRs
  • Review Action Groups
  • Document the incident
  • Create corrective actions

Important Azure Monitor KQL Examples

Find recent failed Azure operations

AzureActivity
| where TimeGenerated > ago(24h)
| where ActivityStatusValue == "Failed"
| project TimeGenerated, Caller, OperationNameValue, ResourceGroup, ResourceId
| order by TimeGenerated desc

Count failed operations by caller

AzureActivity
| where TimeGenerated > ago(24h)
| where ActivityStatusValue == "Failed"
| summarize FailedOperations=count() by Caller
| order by FailedOperations desc

Find recent activity for a resource

AzureActivity
| where TimeGenerated > ago(24h)
| where ResourceId contains "my-vm"
| project TimeGenerated, Caller, OperationNameValue, ActivityStatusValue
| order by TimeGenerated desc

Check VM heartbeat

Heartbeat
| summarize LastHeartbeat=max(TimeGenerated) by Computer
| order by LastHeartbeat asc

This can help identify systems whose heartbeat telemetry is no longer arriving.


Investigate CPU telemetry

Where CPU data is available in InsightsMetrics:

InsightsMetrics
| where TimeGenerated > ago(1h)
| where Name == "Percentage CPU"
| summarize AverageCPU=avg(Val) by Computer, bin(TimeGenerated, 5m)
| order by TimeGenerated asc

Find application exceptions

AppExceptions
| where TimeGenerated > ago(1h)
| summarize ExceptionCount=count() by ExceptionType
| order by ExceptionCount desc

Find failed application requests

AppRequests
| where TimeGenerated > ago(1h)
| where Success == false
| summarize Failures=count() by Name
| order by Failures desc

Find slow application requests

AppRequests
| where TimeGenerated > ago(1h)
| where DurationMs > 2000
| project TimeGenerated, Name, DurationMs, ResultCode
| order by DurationMs desc

Useful Azure CLI Commands

List diagnostic settings

az monitor diagnostic-settings list \
  --resource <resource-id>

List available diagnostic categories

az monitor diagnostic-settings categories list \
  --resource <resource-id>

Query Azure Activity Log

az monitor activity-log list \
  --resource-group <resource-group>

Query metrics

az monitor metrics list \
  --resource <resource-id> \
  --metric "Percentage CPU"

List Action Groups

az monitor action-group list

List metric alerts

az monitor metrics alert list

List Activity Log alerts

az monitor activity-log alert list

List Log Analytics workspaces

az monitor log-analytics workspace list

List Advisor recommendations

az advisor recommendation list

Useful PowerShell Commands

Get Azure Activity Log

Get-AzActivityLog -StartTime (Get-Date).AddHours(-24)

Get Azure metrics

Get-AzMetric -ResourceId "<resource-id>"

List Log Analytics workspaces

Get-AzOperationalInsightsWorkspace

Get Advisor recommendations

Get-AzAdvisorRecommendation

List Action Groups

Get-AzActionGroup

Quick Revision

TopicKey Point
Azure MonitorAzure observability platform
MetricsNumerical time-series data
LogsDetailed event/telemetry records
Activity LogAzure control-plane activity
Resource LogsResource/service-specific telemetry
Log AnalyticsQuery and analyze logs
KQLQuery language for Azure Monitor Logs
DCRDefines data collection and processing
DCRAAssociates a DCR with a resource
AMACurrent Azure Monitor agent
Diagnostic SettingsRoutes supported resource telemetry
Application InsightsApplication performance monitoring
AppRequestsApplication request telemetry
AppDependenciesDependency telemetry
AppExceptionsApplication exception telemetry
Metric AlertAlert based on metrics
Log AlertAlert based on KQL results
Activity Log AlertAlert based on Azure Activity Log
Action GroupDefines alert actions/notifications
Alert Processing RuleControls alert processing
Service HealthAzure platform/service events
Resource HealthHealth of an individual resource
WorkbooksInteractive monitoring reports
AdvisorAzure recommendations
Cost ManagementAnalyze and control Azure spending

Exam Answer Summary

Azure Monitor

Azure’s centralized observability platform for metrics, logs, traces, events and application telemetry.

Log Analytics

The Azure Monitor experience used to query and analyze logs stored in Log Analytics workspaces.

KQL

Kusto Query Language used to query Azure Monitor Logs.

DCR

A Data Collection Rule defines what telemetry is collected, how it is processed and where it is sent.

AMA

Azure Monitor Agent collects supported telemetry and works with DCR-based collection.

Diagnostic Settings

Used to configure supported Azure resources to send platform logs and metrics to destinations such as Log Analytics, Storage or Event Hubs.

Application Insights

Azure Monitor’s application performance monitoring capability.

Metric Alert

Triggers based on metric conditions.

Log Alert

Uses a log query and a condition to trigger an alert.

Action Group

Defines what actions or notifications occur when an alert is triggered.

Service Health

Provides information about Azure service/platform events that may affect your environment.

Resource Health

Shows the health state of a particular Azure resource.

Azure Advisor

Provides recommendations across areas such as reliability, security, performance, cost and operational excellence.


Senior Interview Tip

At senior level, do not answer Azure Monitor questions only by naming portal features.

Explain the observability flow:

Collect
   ↓
Store
   ↓
Query
   ↓
Correlate
   ↓
Alert
   ↓
Respond
   ↓
Improve

For example, if an interviewer asks:

“How would you monitor a production Azure environment?”

A strong answer is:

“I would use Azure Monitor as the central observability platform. I would collect the required resource and application telemetry using appropriate monitoring configurations, use Azure Monitor Agent and DCRs for supported VM data collection, send logs to Log Analytics, use KQL for investigation, Application Insights for application telemetry, metric and log-based alerts for actionable conditions, Action Groups for notifications and automation, Workbooks for operational visibility, and Service Health and Resource Health for Azure platform and resource-level events. I would also regularly review Azure Advisor and monitoring ingestion costs.”

That demonstrates architecture and operational thinking, rather than simply knowing Azure portal menus.


Microsoft Azure Interview Questions & Answers –Series Completed

The Azure section now covers five distinct areas:

Part 1: Azure Networking Fundamentals, VNets, Subnets, NSGs & Connectivity

Part 2: Azure Virtual Machines, Compute, Disks & Availability

Part 3: Azure Storage, Blob Storage, Azure Files, Redundancy & Troubleshooting

Part 4: Load Balancer, Application Gateway, Front Door & Traffic Manager

Part 5: Azure Monitor, Log Analytics, Application Insights, Alerts & Production Monitoring

The five parts together provide a senior-level Azure administration interview foundation without repeatedly asking the same networking, VM, storage or traffic-management questions.


Next Part

Microsoft Intune & Endpoint Management

The next interview series will move to:

  • Microsoft Intune architecture
  • Enrollment
  • Windows Autopilot
  • Configuration profiles
  • Compliance policies
  • Application deployment
  • Windows Update management
  • Device restrictions
  • Endpoint security
  • Defender integration
  • Conditional Access integration
  • Device troubleshooting
  • Hybrid/Azure AD joined device scenarios
  • Real-world enterprise endpoint incidents
  • PowerShell and Intune troubleshooting

The next parts will continue using the same rule: one meaningful question per concept or scenario, with repeated questions avoided.

For official Microsoft 365 documentation and additional technical information, visit: Microsoft Learn

Leave a Comment