Agent Foskett Academy • Microsoft Entra • Module 2 • Lesson 31

Lesson 31 — Microsoft Entra Hybrid Identity Best Practices

Hybrid identity connects on-premises Active Directory with Microsoft Entra ID so users can work across cloud and traditional environments with a consistent identity.

The technology is only one part of the design. Reliable hybrid identity also requires a clearly defined source of authority, an appropriate authentication method, resilient synchronisation, protected administrative accounts, controlled scope, monitored changes and a tested recovery plan.

This lesson brings those decisions together into a practical architecture and operational standard for organisations using Microsoft Entra Connect Sync, Microsoft Entra Cloud Sync or a staged transition between the two.

Hybrid identity should be designed as critical security infrastructure—not as a background synchronisation task that is left unchanged for years.
Agent Foskett Microsoft Entra Hybrid Identity Best Practices lesson
What you will learn

This lesson explains how to design, secure, operate and recover a resilient hybrid identity platform across Active Directory and Microsoft Entra ID.

Source authority and lifecycle
Authentication and resilience
Security and monitoring
Change and recovery planning

Learning objectives

After completing this lesson, you should be able to assess and improve a Microsoft Entra hybrid identity design.

  • Define the source of authority for each identity and attribute.
  • Select an appropriate synchronisation and authentication model.
  • Design high availability without creating conflicting writers.
  • Protect synchronisation infrastructure and privileged access.
  • Control scope, mappings and identity lifecycle changes.
  • Monitor hybrid identity health and security indicators.
  • Plan safe changes, migrations and rollback.
  • Maintain a tested recovery and continuity plan.

The problem this solves

Many hybrid identity environments begin as a small synchronisation project and gradually become critical infrastructure without receiving the same architectural discipline as domain controllers, networks or authentication systems.

Best practices reduce hidden dependencies, conflicting ownership, single points of failure and emergency changes during incidents.

Hybrid identity architecture

Authoritative business process ↓ HR or identity lifecycle system ↓ Active Directory source identity ↓ Controlled scope and attributes ↓ Entra Connect Sync or Cloud Sync ↓ Microsoft Entra ID cloud identity ↓ Authentication + Conditional Access ↓ Applications, devices and resources ↓ Monitoring, governance and response

Principle 1 — Define authority

Every identity and important attribute should have one recognised source of authority. Administrators must know where a name, department, manager, UPN, mail address and enabled state should be changed.

Unclear ownership creates overwritten values, conflicting updates and slow incident response.

Authority questions

  • Who creates the identity?
  • Which system owns employment status?
  • Where are names and departments corrected?
  • Who owns UPN and mail-address standards?
  • Which system disables leavers?
  • Which attributes may be cloud-managed?

Source-of-authority register

Identity propertyRecommended authorityOperational rule
Employment statusHR or approved lifecycle processDisable access from a documented joiner-mover-leaver event
Core account identityOn-premises AD while the object remains synchronisedDo not repair synchronised values only in the cloud
UPN and mail aliasesDocumented identity naming processCheck uniqueness before assignment
Group membershipBusiness owner, role process or governed automationAvoid unmanaged permanent access
Cloud-only propertiesMicrosoft Entra or workload ownerRecord exceptions and prevent competing writers

Principle 2 — Choose the right sync tool

Use Microsoft Entra Cloud Sync when its cloud-managed configuration, lightweight agents and multi-agent resilience fit the organisation's requirements.

Use Microsoft Entra Connect Sync when required features or complex synchronisation scenarios are not supported by Cloud Sync.

Avoid tool selection by habit

The correct decision should be based on supported topology, object types, writeback needs, filtering, transformation, operational model and migration constraints.

Document why the selected platform exists and review the decision as Microsoft Entra capabilities evolve.

Synchronisation design comparison

Design considerationCloud SyncConnect Sync
Configuration modelCloud-managed configuration with lightweight agentsConfiguration and sync engine hosted on a dedicated server
High availabilityMultiple active agents in an agent poolPrimary server with a tested staging-mode server
Operational footprintSmaller local footprintMore local components and database dependencies
Complex scenariosUse when required features are supportedRetain where advanced or unsupported scenarios require it
Change managementManage configuration and agents as production infrastructureManage server, database, rules and upgrades as production infrastructure

Principle 3 — Simplify authentication

Select the least complex authentication method that satisfies business and regulatory requirements.

Password hash synchronization provides cloud authentication and strong service continuity because Microsoft Entra can authenticate users without contacting the on-premises environment during each sign-in.

Authentication dependencies

Pass-through Authentication validates passwords against on-premises Active Directory and therefore requires healthy authentication agents and connectivity.

Federation introduces additional infrastructure and should be retained only where there is a clear requirement that cannot be met with cloud authentication.

Authentication decision guide

MethodOperational characteristicBest-practice consideration
Password hash synchronizationAuthentication occurs in Microsoft Entra IDPrefer for simplicity, scalability and cloud continuity where suitable
Pass-through AuthenticationPassword validation depends on on-premises agentsDeploy redundant agents and test loss of local connectivity
FederationAuthentication depends on federation infrastructureUse only for documented requirements and maintain a cloud-authentication migration plan

Principle 4 — Design for failure

Hybrid identity must continue operating when one server, agent, network path or administrator is unavailable.

Cloud Sync supports multiple active agents. Connect Sync environments should maintain a tested staging-mode server and a documented activation process.

Resilience is more than quantity

Redundant components should not share every dependency. Place agents or servers so one maintenance event, proxy failure or host problem does not remove the entire service.

Test failover rather than assuming that a second installation is usable.

High-availability checklist

AreaBest practiceEvidence
Cloud Sync agentsDeploy multiple active agents on separate supported hostsEach agent independently reports healthy and processes work
Connect SyncMaintain a current staging-mode serverDocumented and tested activation procedure
AuthenticationRemove avoidable on-premises sign-in dependenciesBusiness-continuity test results
NetworkProvide resilient DNS, proxy and outbound connectivityEndpoint tests from every agent host
OperationsEnsure more than one trained administrator can recover the serviceRunbook, access and exercise records

Principle 5 — Protect the sync tier

Synchronisation servers and agents bridge two identity systems and should be treated as sensitive infrastructure.

Use dedicated supported hosts, restrict interactive access, minimise installed software and monitor administrative activity.

Administrative protection

  • Use separate privileged accounts.
  • Require MFA for cloud administration.
  • Use Privileged Identity Management where available.
  • Restrict local administrator membership.
  • Protect service credentials and secrets.
  • Record emergency access procedures.

Security boundaries

Standard user workstation ✕ No administration of sync infrastructure Privileged admin workstation ↓ Time-bound privileged role ↓ Approved maintenance task ↓ Dedicated sync server or agent host ↓ Logged and reviewed change

Principle 6 — Control scope

Synchronise only the objects and attributes required by the approved design. Broad scope increases the blast radius of mistakes and can expose unnecessary directory information.

OU, group and attribute filters should have an owner, purpose and review date.

Scope-change discipline

  • Measure expected object impact.
  • Identify creates, updates and removals.
  • Test with representative identities.
  • Review privileged and service accounts.
  • Record rollback steps.
  • Monitor the first processing cycles.

Safe scope-change workflow

Proposed scope change ↓ Document business requirement ↓ Estimate affected objects ↓ Test with a limited population ↓ Review create, update and delete impact ↓ Approve change and rollback ↓ Implement during monitored window ↓ Verify Provisioning Logs and target objects

Principle 7 — Use stable identifiers

Display names, departments and email addresses can change. Matching and investigation should rely on stable, unique identifiers wherever the supported design permits.

Duplicate or recycled values can create incorrect joins, duplicate cloud identities and difficult recoveries.

Naming and uniqueness

  • Define a UPN naming standard.
  • Validate mail and proxy-address uniqueness.
  • Plan rename and rehire scenarios.
  • Check active and deleted cloud objects.
  • Avoid manual cloud repairs that conflict with sync.
  • Retain object identifiers in incident evidence.

Principle 8 — Minimise customisation

Custom rules and transformations should exist only where a documented business requirement cannot be met by standard behaviour.

Every customisation increases testing, upgrade and troubleshooting complexity.

Customisation register

For each custom rule, record its owner, purpose, source and target attributes, test cases, dependencies, approval date and rollback method.

Remove obsolete rules rather than allowing years of undocumented logic to accumulate.

Change-risk comparison

ChangeRiskRequired control
Add one mapped attributeIncorrect or excessive target valuesTest objects and modified-property review
Change matching logicDuplicate or incorrectly joined identitiesPre-change object comparison and rollback
Expand OU scopeLarge unexpected create or delete impactObject count and privileged-account review
Change authentication methodOrganisation-wide sign-in disruptionStaged rollout and continuity plan
Replace sync platformCompeting writers or missing attributesDocumented migration sequence and validation

Principle 9 — Monitor continuously

Monitor agent health, configuration status, provisioning failures, skipped objects, duplicate values and unusual bulk changes.

Healthy infrastructure should produce expected recent activity—not merely show a green status at one point in time.

Correlate evidence

  • Provisioning Logs
  • Microsoft Entra Audit Logs
  • Sign-in Logs
  • Connect Health or agent status
  • Windows and service logs
  • Source Active Directory changes

Monitoring baseline

SignalBaseline questionAlert condition
Recent activityHow often should each job produce events?No activity beyond the expected interval
Failure rateWhat is the normal number of failed objects?Sharp increase or repeated common error
Skipped objectsWhich skips are expected by design?New skip reason or sudden volume change
Agent healthHow many independent agents should be active?Loss of redundancy or unhealthy agent
Bulk operationsWhat create, update and delete volumes are normal?Unexpected mass lifecycle activity

Principle 10 — Govern the lifecycle

Joiner, mover and leaver processes must operate consistently across Active Directory and Microsoft Entra ID.

Disabling a source account, removing access, revoking sessions and managing retained data are related but separate controls.

Lifecycle evidence

  • Approved employment event
  • Source account creation or disablement
  • Provisioning event and target result
  • Group and role changes
  • Session or token revocation
  • Ticket closure and owner approval

Joiner-mover-leaver control path

Authorised HR event ↓ Identity lifecycle workflow ↓ Source account and attributes ↓ Controlled synchronisation ↓ Cloud identity and access assignment ↓ Governance review ↓ Timely disablement and access removal ↓ Verified logs and retained evidence

Principle 11 — Use staged change

Authentication and synchronisation migrations should use controlled populations, measurable success criteria and a rollback plan.

A staged rollout allows the organisation to validate cloud authentication before changing the whole domain.

Migration controls

  • Inventory current dependencies.
  • Confirm feature compatibility.
  • Select representative pilot users.
  • Synchronise required password data first.
  • Measure sign-in and provisioning outcomes.
  • Remove old infrastructure only after validation.

Migration sequence

Current-state assessment ↓ Target architecture approved ↓ Dependencies and unsupported features identified ↓ Pilot synchronisation or authentication group ↓ Provisioning and sign-in validation ↓ Controlled expansion ↓ Full production transition ↓ Legacy component retirement ↓ Post-migration monitoring

Principle 12 — Prepare recovery

Document how to restore synchronisation, activate standby capacity, replace an agent, recover credentials and validate the target after an incident.

Recovery documentation should be usable by another authorised administrator under pressure.

Recovery package

  • Architecture and data-flow diagram
  • Server, agent and configuration inventory
  • Secure credential recovery process
  • Standby or replacement procedure
  • Known custom rules and dependencies
  • Post-recovery validation checklist

Agent Foskett investigation: “The second server was not a backup…”

1. The production Connect Sync server fails ↓ 2. The team points to a second server labelled “standby” ↓ 3. The server has not been updated for eighteen months ↓ 4. Its configuration no longer matches production ↓ 5. No administrator has tested activation ↓ 6. Agent Foskett compares rules, versions and connectors ↓ 7. The supposed standby cannot be safely activated ↓ 8. Synchronisation is restored from the production recovery plan ↓ 9. A new staging-mode server is built and validated ↓ 10. Quarterly failover testing is added to operations
A second server is not a recovery solution until its configuration, access, activation and target behaviour have been tested.

Security indicators

  • Unexpected changes to synchronisation scope or rules.
  • New privileged accounts entering scope.
  • Unapproved agent or connector registration.
  • Bulk mail, UPN or ownership changes.
  • Administrative access from standard workstations.
  • Removal of monitoring or loss of log activity.

Operational indicators

  • Only one functioning agent or sync server.
  • No tested failover procedure.
  • Undocumented custom rules.
  • Stale or unsupported components.
  • Recurring duplicate-value failures.
  • Unclear ownership for source attributes.

Hybrid identity review checklist

Review areaQuestionEvidence
AuthorityDoes every important identity property have one owner?Source-of-authority register
ArchitectureIs the selected sync and authentication model still appropriate?Current design decision record
ResilienceCan service continue after one component fails?Successful failover exercise
SecurityAre privileged access and sync hosts protected?Role, host and access review
MonitoringWould the team detect stalled or abnormal activity?Alerts, dashboards and incident tests
RecoveryCan another administrator execute the runbook?Documented recovery exercise

Common mistakes

  • Treating synchronisation as a set-and-forget service.
  • Allowing multiple systems to own the same attribute.
  • Keeping federation without a current requirement.
  • Calling an untested server a standby.
  • Synchronising more objects and attributes than necessary.
  • Using permanent privileged access for routine support.
  • Retiring old infrastructure before proving the replacement.

Best practices

  • Prefer simple, supportable architecture.
  • Document authority, scope and matching.
  • Use redundant and independently tested components.
  • Protect the synchronisation tier as identity infrastructure.
  • Baseline normal provisioning behaviour.
  • Stage high-impact changes and migrations.
  • Exercise recovery before an outage.

Key takeaways

  • Hybrid identity is critical security and authentication infrastructure.
  • Every identity and attribute needs a clearly defined source of authority.
  • Select Cloud Sync or Connect Sync from requirements—not habit.
  • Use the simplest authentication design that satisfies the organisation's needs.
  • High availability requires tested independent components and trained administrators.
  • Synchronisation hosts, agents, credentials and privileged roles require strong protection.
  • Scope, matching and custom rules should be minimal, documented and reviewed.
  • Continuous monitoring must cover health, failures, skipped objects and unusual bulk activity.
  • Authentication and synchronisation changes should use staged rollout and rollback.
  • A recovery plan is not complete until another administrator has successfully tested it.

Continue learning

Continue through modern Microsoft Entra authentication, or return to the academy roadmap.

Microsoft Entra Hybrid Identity Best Practices

Microsoft Entra hybrid identity best practices include clear source authority, suitable synchronisation and authentication choices, resilient agents or staging servers, least privilege, controlled scope, stable matching, continuous monitoring and tested recovery.

Microsoft Entra Academy Lesson 31 — Hybrid Identity Best Practices

This Agent Foskett lesson explains how to design, secure, operate, migrate and recover hybrid identity across Active Directory and Microsoft Entra ID using practical architecture and governance controls.