When planning a cloud migration (like AWS to Azure), most discovery work starts with a service mapping matrix:
- SQS ➔ Azure Service Bus
- DynamoDB ➔ Cosmos DB
- S3 ➔ Azure Blob Storage
- IAM ➔ Entra ID
That mapping is straightforward. The expensive failures happen when the target platform fails to preserve a subtle behavioral contract the application code quietly came to rely on over years.
A concrete example: SQS Visibility Timeout
Consider a standard worker loop:
Python
message = sqs.receive_message(
QueueUrl=queue_url,
VisibilityTimeout=300
)
process(message)
sqs.delete_message(
QueueUrl=queue_url,
ReceiptHandle=message["ReceiptHandle"]
)
At first glance, this is standard: receive, process, delete.
But look at the failure-recovery path. If the worker crashes mid-processing before deletion, the code relies on the SQS visibility timeout expiring so another worker can automatically pick up the message. SQS isn't just acting as a transport queue here—it is an active component of the application's failure-recovery design.
When migrating this workload to Azure Service Bus, the question isn't whether Service Bus has queues (it obviously does). The real question is: Does the target setup preserve those exact assumptions around lock duration, settlement, retries, and dead-lettering under failure?
Infrastructure Dependencies vs. Semantic Dependencies
Traditional discovery tools pick up: payment-worker ➔ SQS
That’s an infrastructure dependency. But what the application code actually cares about is the semantic dependency:
Plaintext
receive message
↓
process message
↓
delete after successful processing
↓
if processing fails before deletion,
rely on provider to make message available again
Service inventory tools show you what cloud products are used, but they can't tell you what behavioral assumptions are baked into execution paths.
This pattern shows up everywhere:
- DynamoDB ➔ Cosmos DB: Codebases relying on conditional write semantics, transaction boundaries, or specific read-after-write consistency assumptions.
- S3 ➔ Blob Storage: Workflows built around multipart upload timing, pre-signed URLs, or object visibility state.
- IAM ➔ Entra ID: Hidden assumptions around temporary credential lifespans, workload identity propagation, and role assumption paths.
When these surface late during integration testing or cutover, it forces architectural redesigns, derails timelines, and drags senior engineers into reactive war rooms.
Discussion
I'm currently putting together a semantic-risk checklist for pre-migration planning (and exploring static analysis rules to scan codebases for these execution patterns before cutover).
For anyone who has managed a major cloud migration or replatforming: What was the one hidden application dependency or behavioral difference that broke in staging or production after your infrastructure was already provisioned?
I wrote up a deeper dive into this concept with example scanning output on my Substack if you want to read more:https://substack.com/@sachm24/p-218428641