Packer Template — Unreal Engine Horde Windows Build Agent¶
This template builds a Windows Server 2022 AMI for the Unreal Engine Horde
build/sync agent pools (the samples/unreal-horde-build-pipeline iSCSI/NTFS
thin-clone pipeline). It is a fork of the Jenkins-oriented template in
../windows, adapted for Horde and iSCSI-backed workspaces.
Usage¶
Build in a private / ingress-limited subnet (defaults keep the temporary build instance off the public internet):
packer init .
packer build \
-var "region=us-east-1" \
-var "vpc_id=vpc-xxxxxxxx" \
-var "subnet_id=subnet-xxxxxxxx" \
-var "public_key=<include public ssh key here>" \
windows-horde.pkr.hcl
Or use the provided example.pkrvars.hcl as a starting point:
packer build -var-file="example.pkrvars.hcl" windows-horde.pkr.hcl
The resulting AMI name is prefixed windows-horde-build-agent-<timestamp>. The
Terraform sample looks it up by that prefix via data.aws_ami, so no AMI ID is
hardcoded.
In-VPC build (WinRM stays private) — the real approach used¶
Because associate_public_ip_address = false and ssh_interface = "private_ip",
Packer must be able to reach the temporary Windows instance's WinRM (5986) over
the instance's private IP. That means Packer must run from inside the same
VPC. If your workstation's egress blocks outbound WinRM 5986 to public IPs (or
the build subnet has no public path at all), run Packer from a throwaway Linux
builder launched in the same private subnet (NAT egress for the Windows
instance's Chocolatey/VS downloads; drive the builder via SSM Run Command).
A reference AMI was built this way (the identifiers below are placeholders — substitute your own):
- Throwaway
t3.mediumAmazon Linux 2023 builder in a private subnet (subnet-xxxxxxxxinvpc-xxxxxxxx, e.g.10.0.2.0/24, a single AZ), no public IP, instance profile =AmazonSSMManagedInstanceCore+ a scoped inline EC2/S3 policy sufficient for theamazon-ebsbuilder.packer gitinstalled via the HashiCorp dnf repo. Template staged builder-side via a throwaway S3 bucket.- On the builder (
export HOME=/home/ec2-user):
packer init .
packer build -var-file=build.pkrvars.hcl windows-horde.pkr.hcl
with build.pkrvars.hcl:
region = "us-east-1"
vpc_id = "vpc-xxxxxxxx"
subnet_id = "subnet-xxxxxxxx" # same private subnet as the builder
associate_public_ip_address = false
ssh_interface = "private_ip"
# Packer creates its own temporary SG; scope WinRM 5986 ingress to the
# builder's subnet CIDR. NEVER 0.0.0.0/0.
temporary_security_group_source_cidrs = ["10.0.2.0/24"]
instance_type = "c6a.4xlarge"
root_volume_size = 256
public_key = "<ephemeral ssh public key>"
Build time ≈ 40 min (VS2022 Build Tools dominates). The build fails unless
validate_image.ps1 prints VALIDATION PASSED.
- Tear down all scaffolding (builder instance, instance profile + role, ephemeral keypair, staging bucket, builder SG). Keep only the AMI + snapshot.
Security¶
associate_public_ip_addressdefaults to false andssh_interfacetoprivate_ip— run Packer from inside the VPC (CodeBuild/bastion).- The build never opens WinRM (5986) to
0.0.0.0/0. When nosecurity_group_idis supplied, Packer scopes the temporary SG's ingress to the builder's own IP. If you enable a public IP, you must supply a CIDR-scopedsecurity_group_id.
Installed Tooling (what this AMI bakes)¶
- Chocolatey package manager
- Git
- OpenSSH Server
- Python 3 + botocore + boto3
- Windows Development Kit / Debugging Tools (PDBCOPY)
- Visual Studio 2022 Build Tools
- VCTools workload (include recommended)
- ManagedDesktopBuildTools workload (include recommended)
- MSVC v143 — VC 14.38.17.8 x86/x64 build tools
- Microsoft.Net.Component.4.6.2.TargetingPack
- .NET 6.0 runtime (matches the Horde module
agent_dotnet_runtime_versiondefault =6.0, so the module's first-bootchoco install dotnet-6.0-runtimeis a no-op) - .NET 8 SDK (Unreal Engine 5.5 UnrealAutomationTool)
- Perforce
p4client (matches the module's first-boot install) - AWS CLI (used by the iSCSI/SAN pipeline scripts)
- MSiSCSI initiator service set to Automatic
- Multipath I/O (MPIO) feature enabled + MSDSM set to auto-claim iSCSI (configured after a reboot and verified, or the bake fails)
- Baked ONSTART Scheduled Task (
Horde-SetUniqueIqn) that derives a unique, instance-id-based initiator IQN on every boot (input-free)
Diff vs. the Jenkins template (../windows)¶
| Area | Jenkins template | This (Horde) template |
|---|---|---|
| Jenkins local user | created (setup_jenkins_agent.ps1) |
dropped |
| OpenJDK | installed (Jenkins agent) | dropped |
| NFS-Client | Install-WindowsFeature NFS-Client |
removed (iSCSI instead) |
| .NET 6 runtime | — | added |
| .NET 8 SDK | — | added |
| p4 client | — | added (baked; module no-op) |
| AWS CLI | — | added |
| MSiSCSI initiator | — | added (Automatic) |
| MPIO + MSDSM claim | — | added |
| VS2022 Build Tools + WDK | kept | kept (identical recipe) |
| choco / git / OpenSSH / Python | kept | kept |
| WinRM userdata bootstrap | kept | kept |
| Public IP / interface | public by default | private by default |
| In-build validation | — | added (validate_image.ps1) |
iSCSI baking — the SPIKE (RESOLVED)¶
Question: can MSiSCSI enablement + initiator IQN materialization be fully baked into the AMI, or is a per-boot SSM step unavoidable?
Answer: fully baked. No SSM association is needed. The only true per-boot
iSCSI need is LOCAL and INPUT-FREE: (1) MSiSCSI running, and (2) a UNIQUE
initiator IQN per agent. Both are now baked into the AMI — (1) as service state,
(2) as a baked ONSTART Scheduled Task that derives the IQN from the instance-id
every boot. All ONTAP contact (discovery, login, igroup add, LUN map, mount) and
all Perforce work happen JUST-IN-TIME at JOB time in
buildgraph/attach-clone-lun.ps1 + buildgraph/hydrate-source-lun.ps1, so boot
needs ZERO ONTAP config and ZERO Perforce env.
What is baked (image-time, machine-independent):
- MSiSCSI service → Automatic start (initiator running on first boot, no per-boot enable).
- MPIO feature enabled.
- MSDSM set to auto-claim iSCSI devices — after a reboot, in a separate
provisioner (
configure_mpio_claim.ps1) that verifies the claim and fails the build if it did not stick. The job-time scripts only connect a second iSCSI portal when this claim is active. - A baked ONSTART Scheduled Task (
Horde-SetUniqueIqn, SYSTEM, RunLevel Highest) that runsC:\ProgramData\horde\set_unique_iqn.ps1on every boot.
The initiator IQN — mechanism baked, value per-instance:
On Windows the IQN is stored under
HKLM\SYSTEM\CurrentControlSet\Control\Class\{iSCSI}\...\NodeName and is derived
per install. Baking a fixed IQN into the AMI would give every agent the same
IQN, which breaks the per-agent igroup model on FSxN/ONTAP: each agent must
present a unique IQN so its clone LUN maps to exactly one host, and a collision is
a silent NTFS-corruption trap.
The resolution bakes the mechanism (the ONSTART task + the script) while
leaving the value per-instance. On every boot, set_unique_iqn.ps1:
- reads the EC2 instance-id from IMDSv2 (PUT token, then GET
/latest/meta-data/instance-id) with a bounded retry. The instance-id is mandatory: if IMDS is unavailable the script fails loudly instead of emitting a non-attributable IQN, because the pipeline's igroup self-heal keys on the instance id embedded in the IQN; - sets a deterministic, host-unique IQN
iqn.1991-05.com.microsoft:<instance-id>viaSet-InitiatorPort -NodeAddress <current> -NewNodeAddress <desired>, idempotently (only if different); - ensures MSiSCSI is Automatic + running;
- reads the IQN back, logs it to
C:\ProgramData\horde\set_unique_iqn.log, and fails if it does not match the intended value.
This is input-free: no ONTAP, no Perforce, no Terraform values are needed at boot. Target portal discovery + LUN login stay at job time (targets are only known then).
Recommendation → (a) full baked boot, NO SSM. The sample-side SSM
associations that previously carried a "slim per-boot step" (machine
P4PORT/P4USER + initiator setup) are deleted: the job scripts authenticate to
Perforce explicitly (hydrate-source-lun.ps1 sets its own P4PORT/P4USER from
mandatory params and mints a ticket from Secrets Manager; BuildPipeline.xml
passes p4.exe -p/-u/-c explicitly), and the unique-IQN + MSiSCSI needs are now
baked. Boot is input-free; no modules/ change is required.
In-build validation¶
validate_image.ps1 fails the Packer build (non-zero exit) if any required
component is missing. In addition to the MSVC toolchain, MSiSCSI (Automatic), and
MPIO checks (feature present and MSDSM claiming iSCSI devices), it asserts the
boot-time IQN mechanism: the C:\ProgramData\horde\set_unique_iqn.ps1 file exists
AND the Horde-SetUniqueIqn Scheduled Task is registered — and then runs the
script once and requires the initiator IQN to come back instance-id-derived
(iqn.1991-05.com.microsoft:i-...). Presence of the task is not proof it works:
an earlier revision called Set-InitiatorPort without the mandatory
-NewNodeAddress, threw on every boot, and left each agent on its
hostname-derived default IQN.
Files¶
| File | Purpose |
|---|---|
windows-horde.pkr.hcl |
Packer template (source + build + provisioners) |
userdata.ps1 |
WinRM bootstrap for the Packer communicator |
base_setup.ps1 |
choco, git, OpenSSH, Python (NFS-Client removed) |
install_vs_tools.ps1 |
VS2022 Build Tools + VC 14.38 + WDK/PDBCOPY |
install_horde_agent_tools.ps1 |
.NET 6 runtime, .NET 8 SDK, p4, awscli |
install_iscsi.ps1 |
MSiSCSI Automatic + MPIO feature (the MSDSM claim comes after the reboot) |
configure_mpio_claim.ps1 |
post-reboot: MSDSM iSCSI auto-claim + RR policy, verified (fails the build otherwise) |
set_unique_iqn.ps1 |
baked per-boot script: derives a unique instance-id IQN (dropped at C:\ProgramData\horde\) |
register_iqn_task.ps1 |
image-time step: registers the ONSTART Scheduled Task that runs set_unique_iqn.ps1 every boot |
validate_image.ps1 |
in-build assertion of toolchain + iSCSI + MPIO + MSDSM claim + boot-IQN task (exercised, not just present) |
example.pkrvars.hcl |
example variables (generic, no account values) |