Agent Foskett Academy • Microsoft Entra • Module 2 • Lesson 30

Lesson 30 — Microsoft Entra Cloud Sync Troubleshooting

Microsoft Entra Cloud Sync is designed to provide lightweight, resilient hybrid identity synchronisation, but successful operation still depends on healthy provisioning agents, working network paths, correct scope, valid source data and stable target processing.

When synchronisation stalls, the visible symptom may be a missing user, stale attribute, duplicate object, growing failure count or a configuration placed into quarantine. The fastest route to recovery is a structured investigation that separates agent health, connectivity, scope, matching, attribute and service problems.

This lesson explains how to troubleshoot Microsoft Entra Cloud Sync from the source directory through the agent pool and provisioning service to Microsoft Entra ID, using Provisioning Logs, agent status, object identifiers, error patterns and controlled retesting.

Troubleshoot from evidence outward: confirm the affected objects, establish the failure boundary, then correct and retest the smallest possible component.
Agent Foskett Microsoft Entra Cloud Sync Troubleshooting lesson
What you will learn

This lesson explains how to troubleshoot Microsoft Entra Cloud Sync across provisioning agents, network connectivity, scope, matching, mappings and configuration health.

Agent and service health
Connectivity and agent pools
Scope, matching and mappings
Quarantine and recovery

Learning objectives

After completing this lesson, you should be able to investigate Microsoft Entra Cloud Sync problems methodically and restore synchronisation without relying on trial and error.

  • Define the major Cloud Sync troubleshooting boundaries.
  • Validate provisioning agent and agent-pool health.
  • Investigate service, registration and connectivity failures.
  • Use Provisioning Logs to isolate object-level problems.
  • Diagnose scope, matching and attribute-mapping issues.
  • Recognise quarantine, retry and stalled-processing patterns.
  • Separate single-object failures from configuration-wide outages.
  • Build a safe recovery and verification workflow.

The problem this solves

Cloud Sync incidents rarely present as one clear failure. An agent may appear active while one configuration is unhealthy, one organisational unit is excluded or a target object is blocked by a duplicate value.

A structured workflow prevents unnecessary agent reinstalls, broad scope changes and repeated restarts that hide the original evidence.

Cloud Sync troubleshooting boundaries

Source Active Directory ↓ Object and attribute validity ↓ Scope and filtering ↓ Provisioning agent ↓ Agent pool and registration ↓ Network and service connectivity ↓ Microsoft Entra provisioning service ↓ Matching and attribute mappings ↓ Microsoft Entra target object ↓ Verified synchronisation result

Start with the symptom

Write down exactly what is wrong before changing anything. The troubleshooting path for one missing user differs from the path for an entire configuration that has stopped processing.

  • One object missing
  • One attribute stale
  • Many objects failing
  • No recent provisioning activity
  • Duplicate objects created
  • Configuration in quarantine

Establish the blast radius

Determine whether the issue affects one identity, one organisational unit, one agent, one configuration or every Cloud Sync workflow.

The blast radius is one of the strongest clues to the failure boundary.

Symptom-to-boundary guide

SymptomLikely boundaryFirst evidence
One user missingScope, source data, matching or object-level errorProvisioning Logs for that identity
One OU missingConfiguration scope or directory permissionsScope configuration and source placement
All users in one job failingConfiguration, credentials, mapping or target serviceFailure grouping by job and error
No activity from one serverAgent service, registration, proxy or networkAgent status and local service logs
Duplicate cloud usersMatching rules or non-unique attributesSource and target object identifiers
Configuration quarantinedPersistent high failure rate or unrecoverable processing issueConfiguration health and repeated error pattern

Agent health

Confirm that each provisioning agent is installed, registered and reporting as expected. A healthy agent pool should not depend on a single server.

Check whether the agent service is running, whether the server was recently restarted and whether security software or patching changed local behaviour.

Agent pool resilience

Multiple agents provide resilience only when they are independently healthy and able to reach the required services.

Two agents behind the same failed proxy or network path do not provide meaningful redundancy.

Agent validation checklist

CheckQuestionEvidence
Service stateIs the provisioning agent service running?Windows service status and recent restarts
RegistrationIs the agent associated with the expected tenant and configuration?Portal status and registration details
VersionAre agents running supported and consistent versions?Installed version inventory
Host healthIs the server resource constrained or unstable?CPU, memory, disk, event logs and uptime
Security controlsDid antivirus, firewall or hardening block the service?Protection history and policy changes
RedundancyCan another agent process work successfully?Agent-pool status and recent activity

Connectivity failures

The provisioning agent requires outbound connectivity to Microsoft cloud services and reliable access to the source Active Directory environment.

Proxy authentication, TLS inspection, DNS, firewall rules and service-account restrictions can interrupt processing even when the server remains online.

Connectivity questions

  • Can the agent resolve required endpoints?
  • Can it establish outbound HTTPS sessions?
  • Was a proxy, firewall or inspection policy changed?
  • Can it reach the required domain controllers?
  • Are certificates and system time valid?
  • Is the same failure present on every agent?

Connectivity decision tree

Agent not processing ↓ Is the local service running? ┌────────────┴────────────┐ ↓ ↓ No Yes ↓ ↓ Start service and Check registration inspect local logs and cloud status ↓ Can agent reach cloud services? ┌──────────┴──────────┐ ↓ ↓ No Yes ↓ ↓ Check DNS, proxy, Check source AD access, TLS and firewall scope and object processing

Provisioning Logs

Provisioning Logs remain the primary source for object-level Cloud Sync evidence. Search by identity, status, action, job and time range.

Read the detailed record, not only the summary row. The skip reason, processing step, source and target identifiers and modified properties usually reveal the next action.

Group failures by pattern

Do not investigate hundreds of events individually when they share the same error.

Group by job, status, error code, time and affected object type to identify whether the root cause is common.

Object-level outcomes

OutcomeMeaningTroubleshooting action
SuccessThe recorded action completedVerify target attributes and object state
FailureThe service attempted the operation but could not complete itRead the full error and identify the failing boundary
SkippedThe object was evaluated but no target action was takenReview scope, eligibility, assignment and matching
Repeated retryThe service continues attempting recoveryDetermine whether the issue is transient or persistent
No eventThe object may not have been discovered or the job may not be processingCheck source visibility, scope and agent activity

Scope problems

Cloud Sync can operate normally while expected users remain absent because they are outside the configured organisational units, groups or attribute filters.

Always compare the object's current source placement with the intended scope.

Scope checks

  • Correct organisational unit
  • Required group membership
  • Expected attribute-filter values
  • Configuration enabled and assigned
  • No exclusion rule taking precedence
  • Source object visible to the agent

Missing-object workflow

Expected identity is missing ↓ Search Provisioning Logs ↓ Is there an event? ┌────────────┴────────────┐ ↓ ↓ Yes No ↓ ↓ Read status and reason Check source discovery, ↓ OU scope and agent activity Skipped, failed or success? ↓ Correct source data, matching, mapping or target issue ↓ Re-test one object ↓ Verify Microsoft Entra result

Object matching failures

Cloud Sync must identify whether a source object should join an existing cloud identity or create a new one.

Non-unique or inconsistent matching values can cause duplicate objects, incorrect joins or provisioning failures.

Matching evidence

  • Source object ID
  • Target object ID
  • User principal name
  • Mail and proxy addresses
  • Matching attribute values
  • Soft-deleted or conflicting targets

Matching failure patterns

PatternPossible causeValidation
Duplicate cloud objectExisting target did not satisfy the matching ruleCompare source and target matching values
Duplicate-value errorUPN, mail or proxy address already existsSearch active and deleted target objects
Wrong object updatedMatching value was not uniqueReview stable object IDs and authoritative data
Join never occursSource or target value differs in format or contentCompare exact values, including hidden characters

Attribute-mapping failures

A user may be created successfully while one or more attributes remain missing, invalid or stale.

Review the source value, mapping or transformation, value sent to the target and the target response.

Common mapping issues

  • Required source value is empty
  • Invalid format or excessive length
  • Transformation produces an unexpected result
  • Attribute is not included in the mapping
  • Target value conflicts with another object
  • Source authority is misunderstood

Common Cloud Sync failures

FailureLikely causeFirst action
Agent unavailableStopped service, host issue, registration or network failureValidate agent service, host and portal status
Authentication or authorisation errorCredential, role or directory permission issueReview recent credential and permission changes
Object outside scopeOU, group or attribute filter exclusionCompare source object with configured scope
Duplicate attributeUPN, mail or proxy address collisionSearch target and deleted objects
Invalid attributeUnsupported value, format or transformationReview modified properties and mapping
No recent eventsJob disabled, agent path broken or source not discoveredCheck configuration health and agent activity
Persistent retriesUnresolved transient-looking errorGroup repeated events and identify common cause

Quarantine

A provisioning configuration can enter quarantine when persistent failures make continued normal processing unsafe or ineffective.

Quarantine is a protective state and a symptom of an unresolved root cause—not the root cause itself.

Quarantine mistakes

  • Restarting without reading repeated errors
  • Assuming the agent itself is broken
  • Clearing the state before correcting the cause
  • Ignoring the percentage of failed objects
  • Making broad scope changes during recovery

Quarantine recovery workflow

Configuration reports quarantine ↓ Capture health state and timestamps ↓ Group recent failures by error ↓ Identify the dominant root cause ↓ Correct credential, scope, mapping, duplicate or connectivity problem ↓ Re-test one or a small set of objects ↓ Confirm successful processing resumes ↓ Verify backlog and failure rate decline ↓ Record cause, action and evidence

Configuration-wide failures

When many objects fail at the same time, prioritise shared dependencies: credentials, permissions, agent connectivity, target availability, mapping changes and configuration scope.

A sudden common start time often points to a recent administrative or infrastructure change.

Single-object failures

When other users continue to synchronise, focus on the affected source object, scope eligibility, unique values, matching and attribute quality.

Do not reinstall agents to solve a problem isolated to one identity.

Safe recovery workflow

Define the symptom ↓ Establish time and blast radius ↓ Check configuration and agent health ↓ Review Provisioning Logs ↓ Identify the failure boundary ↓ Validate source, scope, matching and mappings ↓ Correct one root cause ↓ Re-test a controlled object ↓ Verify target result ↓ Monitor backlog, retries and recurrence ↓ Document evidence and closure

Agent Foskett investigation: “The synchronisation stopped…”

1. New employees stop appearing in Microsoft Entra ID ↓ 2. Both provisioning agents still report active ↓ 3. Existing cloud users can still sign in ↓ 4. Provisioning Logs show repeated failures ↓ 5. The configuration enters quarantine ↓ 6. Agent Foskett groups the failures by error ↓ 7. Every failure references a duplicate proxy address ↓ 8. A bulk HR import reused an old email alias ↓ 9. The conflicting source values are corrected ↓ 10. One affected user is re-tested successfully ↓ 11. Normal processing resumes and the backlog clears ↓ 12. The incident record captures the duplicate-value root cause
The agents were healthy. The service was reachable. A repeated source-data conflict pushed the configuration into quarantine.

Common mistakes

  • Reinstalling agents before checking object-level logs.
  • Assuming active means fully healthy.
  • Ignoring skipped events.
  • Changing multiple settings at once.
  • Clearing quarantine without correcting the cause.
  • Testing recovery with the entire directory.
  • Closing the incident before verifying target objects.

Best practices

  • Maintain at least two independently healthy agents.
  • Baseline normal event and failure volumes.
  • Keep scope narrow and documented.
  • Use stable, unique matching attributes.
  • Validate mapping changes with test objects.
  • Capture identifiers and timestamps for every incident.
  • Document root cause, correction and verification.

Key takeaways

  • Cloud Sync troubleshooting should begin with the exact symptom and blast radius.
  • Agent health, connectivity, scope, matching, mappings and target processing are separate failure boundaries.
  • A healthy agent does not prove that every configuration or object is synchronising.
  • Provisioning Logs provide the strongest object-level evidence.
  • One missing user usually requires a different investigation from a configuration-wide outage.
  • Skipped events commonly reveal scope and eligibility problems.
  • Duplicate attributes and matching failures can create widespread provisioning errors.
  • Quarantine is a protective state caused by persistent failure, not a diagnosis.
  • Recovery should correct one root cause and retest a controlled object.
  • Verification must include target objects, failure trends, agent health and recurrence monitoring.

Continue learning

Continue through Microsoft Entra hybrid identity, or return to the academy roadmap.

Microsoft Entra Cloud Sync Troubleshooting

Microsoft Entra Cloud Sync troubleshooting includes validating provisioning agents, agent pools, connectivity, configuration health, source scope, object matching, attribute mappings, Provisioning Logs, retries and quarantine.

Microsoft Entra Academy Lesson 30 — Cloud Sync Troubleshooting

This Agent Foskett lesson explains how to diagnose missing identities, stalled synchronisation, duplicate objects, persistent failures and quarantined Cloud Sync configurations using a repeatable evidence-driven workflow.