DEV Community

Alessandra Bilardi
Alessandra Bilardi

Posted on Originally published at alessandra.bilardi.net

Managing private data from the web securely on AWS

bookmarks architecture on AWS

From a list of links to a collection of private files

A bookmarks service, in its classic form, is a list of links: a title, an address and a few labels. This time I needed more: MP3s and PDFs, like synthesized podcasts for studying and their handouts, to listen to and read again while keeping track of where I left off, to share with a few people and to keep away from everybody else.

My bookmarks have been public since 1998, and the previous version, bookmarks-v3.0, was a static site generated by Jekyll and published on GitHub Pages: it cost nothing, and for a list of public links it was good enough. The criterion stayed the same, spending as little as possible, but with a new constraint: the data has to stay private. Keeping it private takes two things: pages that ask for the data only after the login, and hosting that protects those pages. v3.0 has neither:

  • GitHub Pages: it is the hosting, and it serves the same files to everybody. It does not let you set HTTP headers: a Content Security Policy (CSP) can only go inside the page, in a reduced form, and it lacks the headers that forbid embedding the page in another site, enforce HTTPS and keep the address you came from away from other sites. And since it serves only static files, it cannot host an API, which would then have to live on another origin
  • Jekyll: it is the generator, and it creates the pages at build time, so a page contains only what was there back then, and private data cannot go into a published build. On top of that, outside GitHub Pages, which runs it by itself, Jekyll would mean keeping Ruby and its gems up to date

I dropped the other options for more mundane reasons. I had two repositories of my own, designed exactly for a static site with its resources, aws-static-website and aws-static-gui-resources: since February 2022 they have been stuck on AWS Cloud Development Kit (CDK) v1, out of support since June 2023, and their bucket is public, which is exactly the opposite of what I needed. Renewing them meant rewriting them.

Six choices, all to keep the data where it belongs

As a lazy developer, I started from what I had already built and what worked: aws-card-clash, the site for running tournaments written with AWS Serverless Application Model (SAM) and TypeScript, which already had a private bucket behind Amazon CloudFront, an authorizer checking the tokens at the entrance of the API, Amazon Cognito federated with Google and a single Amazon DynamoDB table. The difference is that there, unauthorized access would at most let you see a tournament; here it would let you open somebody else's files. Every choice is weighed against that. How the pages work, the commands and the costs are in the README.

CloudFront, without Route 53

The site had to live in a private bucket and send the security headers in the HTTP response. The alternatives were staying on GitHub Pages, with the API on another domain, or moving to CloudFront.

I chose CloudFront in front of a private Amazon Simple Storage Service (S3) bucket. With Origin Access Control only the distribution reads the bucket, and the API answers under /api, on the same origin as the site. So there is no need for Cross-Origin Resource Sharing (CORS) between site and API, the CSP is complete, set in a response headers policy of the distribution, and changing the site takes AWS credentials, not just a push.

It costs the same, which is nothing, because the CloudFront free tier covers a terabyte a month. As for Amazon Route 53, I did not use it: the hosted zone costs 50 cents a month, and two CNAME records written by hand in the DNS I already had do the job, one to validate the AWS Certificate Manager (ACM) certificate, which is free, and one that aliases the name of the site to the distribution.

Invitation only, and short-lived tokens

Whoever gets in has to be somebody I invited. The alternatives were email and password managed by Cognito, or the Google login: I chose Google, because that way I keep nobody's password and, let's be honest, because with aws-card-clash it came for free. The list of invitations lives in DynamoDB, and the one reading it is a pre sign-up trigger: an AWS Lambda function that Cognito calls before creating a user, and that refuses every address without an invitation.

Then I had to decide where to keep the tokens. The sturdier alternative on paper was HttpOnly cookies, which a script in the page cannot read; but a script in the page does not need the token to read the private data: it reads it with the session of the page itself. The cookie would only have prevented copying the token, at the price of writing login, refresh and logout by hand.

I chose sessionStorage with the shortest lifetimes Cognito allows: five minutes for the token and an hour for the refresh. Unlike localStorage, sessionStorage goes away with the tab: closing the page closes the session, and a copied token is good for an hour at most.

The real defense is making it unlikely that a script gets into the page: a CSP that accepts only the code of the application bundle, with no inline code and no eval, texts always inserted as text and never as HTML, and only http and https links.

Monitoring instead of limiting

Whoever uploads files generates costs, and whoever deployed has to know who costs how much. The alternatives were quotas per person, a maximum size, or letting only the curator upload.

I chose to limit nobody and to count: every download counts a request and the bytes of the file, every upload its bytes, and everybody sees their own usage with the costs next to it. The people are few and invited by me, which is why knowing who costs how much is good enough, without writing a quota system to keep running.

As a lazy developer, I did not write the prices in the code: the deploy reads them by itself from the AWS Price List Query API and passes them to the Lambdas. If AWS renames an entry of the price list and a price is not found, the deploy stops instead of showing a cost of zero.

Handling the files

A thirty-minute MP3 weighs tens of megabytes. The alternatives were serving the files from the API, or giving the browser direct access to S3. The API is out: a Lambda answers with 6 MB at most, and even under that limit you would pay for the time of every download.

I chose presigned URLs: the Lambda that handles the files checks who may have the file and signs an address that lasts fifteen minutes for the upload and an hour for the download, enough to listen to a podcast with a few breaks; the browser talks directly to the bucket, which stays private.

Precisely because the file goes from the browser to the bucket without passing through the backend, the backend does not know whether the upload finished, nor how big the file really is: the browser could stop halfway, or declare a fake size. Only S3 knows, and once the upload is done it sends an event with the real size: that event marks the file as ready and counts the uploaded bytes for the monitoring. From then on, every presigned URL for a download counts the size recorded by the event.

From the upload on, the files sit in Amazon S3 Intelligent-Tiering, and the ones nobody opens move down by themselves to a cheaper class, with no lifecycle rules to write.

A public page at a fixed cost

Some links have to be visible to everybody, without a login, and that is the only point of the system open to anyone. The alternatives were a public route of the API served by the same Lambda as the others, or a Lambda of its own.

I chose a Lambda of its own, so its role can only read, and the answer contains the title, the link, the tags and nothing else. Above all, in front of the Lambda there is a CloudFront cache policy, which keeps the answer of that one path for five minutes. A thousand visits in five minutes become one call to Lambda and one DynamoDB read for each CloudFront edge location they come from, whatever the traffic.

The price is that a publication shows up, or goes away, up to five minutes late, which for public links on a personal site is perfectly fine.

The cache, though, protects only whoever goes through CloudFront: the HTTP API of Amazon API Gateway also has an address of its own, reachable directly, and an open route lets every request through to the function. It is the door to a denial of wallet, an attack that aims not to bring the service down, but to run up the bill. At first I handled it with throttling on the route: one request per second, with a burst of five. However, that solution does not close the door.

I closed it by giving the function an address of its own, a function URL protected by Origin Access Control like the bucket of the pages: CloudFront signs every request, and a call that does not come from the distribution gets a 403 error before the function starts. The test shows that it never becomes an invocation, so it is never billed. The route of the HTTP API stays, behind the authorizer like all the others, only for local runs, where there is no function URL.

Handing the data over without leaving it lying around

When a person is removed, their data has to be deleted, but first it has to be handed over to them. The alternatives were sending them the files by hand, or preparing an archive.

I chose a zip on S3, reachable through a presigned URL that lasts seven days, and in the import format, so the person can load it again into a system of their own.

Then a lifecycle rule on the prefix of the exports deletes the zip after seven days, the same lifetime as the link. S3 counts the days from the creation and rounds up to the next midnight UTC: the zip can outlive the link by up to a day. From the expiration of the zip on, nothing is billed, and there is nothing left to remember to delete.

What I did not see coming

Starting from a base that was already thought through, the surprises were few: three, and none of them in the code of the application.

The containers that would not go away

Locally the Lambdas run with sam local, which creates a new container for every call, and a page makes six. With warm containers, which stay up between one call and the next, the automated tests of the API get a lot faster. Except that, once the tests were over, some containers were still up, for two reasons.

The first: the test script, local/api-check.sh, starts sam local in the background with &, and SAM shuts its containers down only when it receives SIGINT, the signal sent by Ctrl+C. But POSIX makes the processes started with & by a script ignore SIGINT, and the signal, even though it arrived, was thrown away. The env command, with --default-signal=INT, puts SIGINT back to its normal behavior before starting SAM, and from there the script shuts it down properly.

The second: by default sam local keeps no container up, and warm containers have to be asked for with --warm-containers, which has two modes. I had picked EAGER, which creates them all at startup yet, with nested stacks, creates a duplicate at the first call, and at shutdown stops only one of the two. LAZY, which creates them at the first call, does not have the problem.

With these two changes, once the tests are over no container is left.

A seven-day link that lasts an afternoon

For the export I had planned a seven-day presigned URL, the maximum S3 allows. Rereading the documentation of presigned URLs before writing the code, I got a surprise: a presigned URL is valid as long as the credentials that signed it are.

With temporary credentials, such as Single Sign-On (SSO), an assumed role or an instance profile, the link expires with the session, after a few hours, even if you ask for seven days. And a Lambda does not help, because its role has temporary credentials too. All things considered, whoever follows best practices ends up with a link that lasts an afternoon.

There are three ways to really get seven days:

  • a dedicated AWS Identity and Access Management (IAM) user: with long-term access keys and the sole permission to read the exports, it is the simplest, but those are exactly the keys best practices advise against, and the ceiling stays at seven days
  • a CloudFront signed URL: the signature does not come from IAM, but from a key pair created for the purpose. The public key is registered on the distribution, the private one stays with whoever signs. Since the signature depends on no session, the expiration can be whatever you want. It costs nothing, because the keys are free, the traffic falls within the CloudFront free tier, and the private key can sit in a standard parameter of AWS Systems Manager Parameter Store, which is free. In exchange, the zip has to be served by CloudFront instead of S3
  • a link that regenerates the link: you send the person the address of an API route with a code, saved with a seven-day expiration, for example with the DynamoDB Time To Live (TTL). At every opening, a Lambda checks that the code has not expired and that the zip still exists, and signs, on the spot, an S3 URL that lasts a few minutes. Temporary credentials are enough, because every signature is fresh

A certificate blocked by GitHub, and a name that existed but could not be found

The site was supposed to live on bookmarks.alessandra.bilardi.net. Once the validation record was written in the DNS panel, the certificate stayed pending. The record was right: the public resolvers of Google and Cloudflare answered with the expected value.

The culprits were the Certification Authority Authorization (CAA) records, which say which authorities may issue a certificate for a name: if the name has none, you climb to the domain above, following the CNAMEs. And alessandra.bilardi.net is a CNAME to my GitHub Pages, whose CAA records allow Let's Encrypt, DigiCert and Sectigo, but not Amazon.

The only place to add a CAA allowing Amazon was the name of the site itself, because alessandra.bilardi.net is a CNAME, and nothing else can be added to it. However, the name of the site had to become a CNAME too, to CloudFront, with the same limit: a CNAME does not live alongside other records, so the CAA would have had to go away exactly when the site went online, and the first renewal, which checks the CAA records again, would have failed.

That is how the site ended up on bookmarks.bilardi.net, where there are no CAA records. One constraint stays for the future: if one day I add CAA records on bilardi.net, for another authority, they will have to allow Amazon too, or the renewal of the bookmarks certificate will fail.

With the certificate issued and the second CNAME written, the checks still failed, with Could not resolve host: the resolver of my computer had asked for the name during the deploy, before the record existed, and kept the answer "no such name" in memory. The site was already answering: forcing the CloudFront address was enough to see it. I have lost count of how many times this has happened to me: resolvectl flush-caches empties the cache, and everything falls into place.

What if it became a service ?

As it is, the system is good enough for me and for the few people I invite. Private data is sensitive anyway, mine or somebody else's; a service, on the other hand, would have to be offered and guaranteed to people I do not know.

What does it take to offer a service ?

  • Terms and conditions: the curator sees everybody's data, to be able to help when something goes wrong. Among people who know each other, saying so in the footer of the pages is enough; for strangers it would take terms to accept, and access limited to when it is needed
  • Quotas: among a few people, counting the usage is enough; in a service it would take a quota per person and a maximum file size, so I do not pay for what somebody else uploads
  • Prices that update themselves: today the prices update only when I deploy again; as a lazy developer, I would let a scheduled Lambda read them again every month from the Price List Query API and keep their history, so that every month is estimated with the rates in force back then
  • The export from the pages: today it is the curator who hands over the data of a person; in a service everybody should be able to download it by themselves, at any time
  • A Command Line Interface (CLI) for whoever uses the service: for example to import their own bookmarks, which today only the curator can do. The API accepts any client with a valid token of the pool, which means a login from the terminal and one more callback address on Cognito would be enough, without touching the backend

What does it take to guarantee a service ?

  • A long cache, always up to date: five minutes is a compromise, because a link taken off the public page by mistake stays visible all that time. With more curators or more traffic, a one-day cache, invalidated every time a link is published or taken away, is the better deal: the first thousand invalidations a month are free, and the Lambda that handles the items needs the permission to make them and the ID of the distribution
  • Dependencies to keep up with: updating them regularly closes the known bugs, but a new version can bring its own, or have been compromised. The latest stable version is the better bet, not the one released the same day: a compromised version, once discovered, gets pulled, and waiting a few days gives it time to happen
  • CloudFront signed cookies: with presigned URLs, a copied link opens the file until it expires; with cookies the link alone opens nothing. They cost a key pair to look after and one more path in the distribution, which costs nothing; the private key can sit in a standard SecureString parameter of Parameter Store, which is free, like the secret of the Google client
  • Intrusion detection: logging the reads of the files prevents nothing, though it lets you find out afterwards who read what; the next step is a Lambda that reads those logs and flags the unusual reads
  • An immediate ban: today whoever is removed keeps the token they already have, for five minutes at most, because the authorizer checks the signature of the token with a public key it keeps in memory and does not ask Cognito whether the user still exists. Checking the invitation at every call would close those five minutes too, at the cost of one table read per request

None of these points requires redoing the architecture: they are all pieces to add to the one that is there, on the day they are needed.

Top comments (0)