Originally published on kuryzhev.cloud
A nightly copy from an on-premises NFS share into S3 works fine for weeks, then stops. Nobody notices until a downstream report is empty, because the task failed quietly and no alarm was wired to it. DataSync EventBridge automation exists to close that gap. It turns transfer state changes into events, and events into schedules, alerts, and follow-up jobs.
This checklist is for engineers who already have a DataSync task running manually and want it to run unattended. It assumes S3 as the destination. Details such as event field names and quotas change over time, so verify them against the DataSync user guide before you rely on them.
Why this checklist
Many unattended DataSync failures are not service problems. They come from missing glue: nothing starts the task on time, nothing tells anyone when it fails, and nothing reacts when the data lands. Each gap is small, and each is easy to close once you know where to look.
A checklist suits this topic better than a tutorial because the pieces are independent. You can set up scheduling, eventing, permissions, and cost controls in any order, and skipping one rarely causes an error on day one. The failure shows up later, usually on a night when something unusual happens. A source mount might go offline, or a task execution might overlap with the previous one.
The pattern covered here has three parts:
- Trigger: EventBridge Scheduler starts the task, or an upstream event does.
- Observe: DataSync task execution state changes flow to EventBridge.
- React: S3 object events or a successful execution kick off processing, and failures page someone.
Even the observe step alone removes a significant blind spot.
The checklist
-
Decide who owns the schedule. DataSync has a built-in task schedule, and EventBridge Scheduler can call
StartTaskExecutiondirectly. Pick one. Running both produces overlapping executions that are hard to explain later. Scheduler gives you time zones and flexible windows, so it is often the better choice. -
Create a dedicated Scheduler role. It needs
datasync:StartTaskExecutionon the specific task ARN and nothing broader. If you configure a dead-letter queue, it also needssqs:SendMessageon that queue. Its trust policy must allowscheduler.amazonaws.com. - Create the schedule with the universal target. The example below uses the AWS SDK target, so you do not need a Lambda function just to start a task.
- Enable EventBridge delivery on the destination bucket. S3 sends events to EventBridge only when the bucket is configured for it.
- Write a failure rule first. Match DataSync task execution state changes that end in an error. Route them to SNS, a chat webhook, or an incident tool.
- Write a success rule second. Use it to trigger post-processing, such as a validation job, a catalog update, or a manifest write.
- Scope S3 object rules by prefix. Filter on bucket name and key prefix so unrelated uploads do not trigger your pipeline.
- Turn on CloudWatch Logs for the task. Set the log level deliberately. Basic is usually enough for routine runs, while transfer-level detail costs more to store.
- Enable task reports to an S3 location if you need per-file evidence of what was transferred, skipped, or failed.
- Handle target failures explicitly. A dead-letter queue and retry policy on a rule or schedule target only catches events that could not be delivered. Throttling and missing permissions are typical causes. Errors inside your own code are a separate case. For Lambda, which EventBridge invokes asynchronously, configure an on-failure destination or a function-level dead-letter queue as well. A target that fails silently is the same blind spot you started with.
First, the failure rule. This event pattern matches DataSync task execution errors. Verify the exact detail-type and field names against the current EventBridge event reference for DataSync.
{
"source": ["aws.datasync"],
"detail-type": ["DataSync Task Execution State Change"],
"detail": {
"State": ["ERROR"]
}
}
Next, the schedule. This CLI call starts a task nightly in a named time zone. It uses the SDK target, so no compute is involved. It also sends failed start attempts to an SQS dead-letter queue, for example when the call is rejected or throttled.
aws scheduler create-schedule \
--name nightly-datasync-nfs-to-s3 \
--schedule-expression "cron(0 2 * * ? *)" \
--schedule-expression-timezone "Europe/Berlin" \
--flexible-time-window Mode=OFF \
--target '{
"Arn": "arn:aws:scheduler:::aws-sdk:datasync:startTaskExecution",
"RoleArn": "arn:aws:iam::111122223333:role/scheduler-datasync-start",
"Input": "{\"TaskArn\":\"arn:aws:datasync:eu-central-1:111122223333:task/task-0example\"}",
"RetryPolicy": {"MaximumRetryAttempts": 2, "MaximumEventAgeInSeconds": 3600},
"DeadLetterConfig": {"Arn": "arn:aws:sqs:eu-central-1:111122223333:scheduler-datasync-dlq"}
}'
# Input is a JSON string inside JSON, so the inner quotes must be escaped
Commonly missed items
These are documented behaviors and typical failure modes that tend to surface after the pipeline has been live for a while.
Watch out for bucket notification overwrites. The S3 API call that enables EventBridge delivery replaces the bucket's whole notification configuration. If the bucket already has Lambda or SQS notifications, sending only {"EventBridgeConfiguration":{}} removes them. Read the current configuration first and merge. Infrastructure-as-code tools can also fight each other over this one setting.
# Read existing config before changing anything
aws s3api get-bucket-notification-configuration --bucket landing-bucket
# Then write back the merged result, including the EventBridge key
aws s3api put-bucket-notification-configuration \
--bucket landing-bucket \
--notification-configuration file://merged-notifications.json
Watch out for event loops. If a rule triggers a job that writes back into the same bucket and prefix, each write emits another object event. Separate the input and output prefixes, or use a different bucket, and make the rule's prefix filter strict.
Other items that regularly get skipped:
- Success does not mean complete data. A task can finish while skipping files, depending on its filters and options. Check the verification settings. Where it matters, compare counts from the task report.
- Overlapping executions. A task generally runs one execution at a time. What happens when a second start is requested depends on current service rules. Check the documented quotas and queuing behavior rather than assuming.
- Service-created objects. DataSync can write helper or metadata objects to destinations. Look at what actually lands in your bucket and exclude those objects from your S3 rules.
- Storage class choice. Writing straight into an archival class carries minimum storage duration charges. If the data is reprocessed soon, that can cost more than a standard class plus a lifecycle rule.
- Cross-account and KMS. A bucket in another account, or one encrypted with a customer managed key, needs grants for the DataSync location role. Grant access in both the bucket policy and the key policy.
- Region mismatch. DataSync and S3 events are delivered to the default event bus in the Region where they occur. The rule that matches them must live in that Region. From there, you can forward events to a bus in another Region if your alerting is centralized.
Automation ideas
Once the basics are stable, a few small additions add value without growing the architecture.
Chain stages with events, not timers. Do not schedule a processing job at 03:00 and hope the transfer finished. Instead, match the success state change and start the next stage from that. This removes the guessing and the fragile time gaps. A Step Functions state machine works well as the target when the follow-up has several steps.
Use input overrides for ad-hoc runs. StartTaskExecution accepts include and exclude filters and option overrides. A second schedule can transfer only a recent prefix, such as the current day's folder, while the nightly run does the full pass. Stagger the two so they do not overlap, because both start executions of the same task. Verify the filter syntax in the DataSync documentation, since patterns are pipe-delimited and easy to get subtly wrong.
Build a freshness check. An event-driven alarm tells you about failures but not about a task that never started. Have the success rule's target publish a custom CloudWatch metric on each successful execution. Then alarm on that metric over your expected interval, with missing data treated as breaching. This catches disabled schedules, broken Scheduler permissions, and offline agents.
Define everything in code. Keep the task, schedule, role, rules, and bucket notification in one Terraform module or CloudFormation stack, so the notification overwrite problem has a single owner. For more AWS automation patterns, browse the rest of kuryzhev.cloud.
Tag and route by environment. Use rule input transformers so one alerting target can name the failed task and execution in its message. In DataSync execution events, the execution ARN appears in the event's resources array rather than in detail, so map it from there. Responders then see which job failed without opening the console. The EventBridge user guide covers input transformers, retry policies, and dead-letter queues in detail.
Start with the failure rule and the freshness check. Together they cover both a failed run and a run that never started, which are two of the easiest ways for this pipeline to fail unnoticed.
Top comments (0)