LaunchKit
← All posts
· 16 min read · by The LaunchKit team · 2 views

Ruby on Rails on AWS Lambda

A Rails app on AWS Lambda fails at config/boot.rb line 3, and the error you are shown is a NoMethodError inside Amazon's own runtime client rather than the thing that actually went wrong. Everything else about this works better than it has any right to: nine lines of handler, a Dockerfile, and a real HTTP response out of a real execution environment. The order matters, so this page goes in the order you will meet it.

Conditions for every measurement below. Apple M2 Max, 12 cores, macOS arm64-darwin25, Docker Desktop 4.37.2, engine 27.4.0. The function is a generated Rails 8.1.3.1 app on pg, built on public.ecr.aws/lambda/ruby:4.0 (Ruby 4.0.7, Amazon Linux 2023, aws_lambda_ric-3.2.0, bundler 4.0.20, json 2.18.0) and run through the AWS Lambda Runtime Interface Emulator that ships inside that image at /usr/local/bin/aws-lambda-rie. PostgreSQL 17.7 on the host. Nothing here was deployed to an AWS account: the numbers are the emulator's own REPORT lines, on this laptop's CPU, and they are not Graviton numbers. Where a figure comes from AWS documentation rather than from a run, the sentence says so and names the page.

The whole handler is nine lines

app.rb sits at the root of the application, which is what Lambda calls /var/task, and it is the entire integration:

require_relative "config/boot"
require "lamby"
require_relative "config/application"
Rails.application.initialize!

def handler(event:, context:)
  Lamby.handler Rails.application, event, context
end

Rails.application.initialize! runs during the Init phase, outside the handler method, which is the one design decision that matters: boot happens once per execution environment and the invocations after it are free of it. lamby 7.0.0 is doing the work in the last line, converting an API Gateway event into a Rack env and the Rack triplet back into the JSON shape Lambda expects.

The Dockerfile is the other half. public.ecr.aws/lambda/ruby:4.0 already contains the runtime interface client and its bootstrap, so there is nothing to install for Lambda itself; gcc, make and libpq-devel are there for pg:

FROM public.ecr.aws/lambda/ruby:4.0

RUN dnf install -y gcc make libpq-devel && dnf clean all

ENV BUNDLE_GEMFILE=/var/task/Gemfile \
    BUNDLE_PATH=/var/task/vendor/bundle \
    BUNDLE_WITHOUT=development:test \
    BUNDLE_FROZEN=true \
    RAILS_ENV=production

WORKDIR /var/task
COPY Gemfile Gemfile.lock ./
RUN bundle install && bundle clean --force

COPY . .
RUN SECRET_KEY_BASE_DUMMY=1 bundle exec rails assets:precompile

CMD [ "app.handler" ]

BUNDLE_PATH=/var/task/vendor/bundle is not arbitrary. Line 10 of the image's /var/runtime/bootstrap reads:

export GEM_PATH=${GEM_PATH:+$GEM_PATH:}/var/task/vendor/bundle/ruby/4.0.0:/opt/ruby/gems/4.0.0:/var/runtime:/var/runtime/ruby/4.0.0:/var/lang/lib/ruby/gems/4.0.0

That path is first, ahead of the runtime's own gems, which is Lambda telling you where a Ruby function's bundle is meant to live.

Run it, POST an API Gateway payload 2.0 event at the emulator, and a Rails response comes back:

$ curl -s -XPOST "http://localhost:9015/2015-03-31/functions/function/invocations" -d @event-hello.json
{"statusCode":200,"headers":{"x-frame-options":"SAMEORIGIN","x-xss-protection":"0","x-content-type-options":"nosniff","x-permitted-cross-domain-policies":"none","referrer-policy":"strict-origin-when-cross-origin","content-type":"text/plain; charset=utf-8","vary":"Accept","etag":"W/\"2f57017a45244136937a32a4f80cef35\"","cache-control":"max-age=0, private, must-revalidate","x-request-id":"699b747d-ee91-4137-a667-64bce8495f4e","x-runtime":"0.008236"},"body":"hello from pid 16 boot 1790524514.763297\n"}

One missing checksum in the lockfile stops the boot

Lambda mounts the function's code read-only and gives you /tmp and nothing else. Reproduce that with docker run --read-only --tmpfs /tmp:rw,size=512m, and on the first attempt here the container never reached Rails:

/var/lang/lib/ruby/4.0.0/bundler/shared_helpers.rb:122:in 'Bundler::SharedHelpers#filesystem_access': There was an error while trying to write to `/var/task/Gemfile.lock`. File system is read-only. (Bundler::ReadOnlyFileSystemError)
    from /var/lang/lib/ruby/4.0.0/bundler/definition.rb:417:in 'Bundler::Definition#write_lock'
    from /var/lang/lib/ruby/4.0.0/bundler/definition.rb:398:in 'Bundler::Definition#lock'
    from /var/lang/lib/ruby/4.0.0/bundler/runtime.rb:104:in 'Bundler::Runtime#lock'
    from /var/lang/lib/ruby/4.0.0/bundler/runtime.rb:35:in 'Bundler::Runtime#setup'
    from /var/lang/lib/ruby/4.0.0/bundler.rb:166:in 'Bundler.setup'
    from /var/lang/lib/ruby/4.0.0/bundler/setup.rb:32:in 'block in <top (required)>'
    from /var/task/config/boot.rb:3:in 'Kernel#require'
    from /var/task/config/boot.rb:3:in '<top (required)>'

Nothing in that trace is Rails. bundler/setup, on line 3 of config/boot.rb, decided the lockfile needed rewriting. The decision is one line of runtime.rb:

def lock(opts = {})
  return if @definition.no_resolve_needed?
  @definition.lock(opts[:preserve_unknown_sections])
end

With a lockfile bundler is happy with, no_resolve_needed? is true, write_lock is never reached, and the same image boots read-only with no special configuration. The lockfile in that first build was not one bundler was happy with. It had been produced by a bundle install that added lamby to an already-resolved lock, and that path writes the gem into CHECKSUMS without its sha256. One entry with no checksum is a resolve, a resolve is a write_lock, and write_lock reaches this:

if File.exist?(file) && lockfiles_equal?(@lockfile_contents, contents, preserve_unknown_sections)
  return if Bundler.frozen_bundle?
  SharedHelpers.filesystem_access(file) { FileUtils.touch(file) }
  return
end

FileUtils.touch on a read-only mount is Errno::EROFS. You can recreate the whole failure from a working image with one substitution, which is how the chain above was confirmed rather than guessed:

RUN sed -i 's|^  lamby (7.0.0) sha256=.*|  lamby (7.0.0)|' /var/task/Gemfile.lock

The dead end. return if Bundler.frozen_bundle? sits right there on the line above the touch, so BUNDLE_FROZEN=true looks like the fix, and it is not. Frozen mode refuses the same lockfile earlier and louder:

/var/lang/lib/ruby/4.0.0/bundler/definition.rb:491:in 'Bundler::Definition#ensure_equivalent_gemfile_and_lockfile': Your lockfile has an empty CHECKSUMS entry for "lamby", but can't be updated because frozen mode is set (Bundler::ProductionError)

bundle lock --add-checksums did not fill the gap either, and the checksums were never actually missing from anywhere but the lockfile: curl https://index.rubygems.org/info/lamby returns checksum:d494f52b... for 7.0.0. What worked was deleting Gemfile.lock and resolving from nothing, in a container on the same base image so the platform-specific gems resolve for aarch64-linux:

  lambda-console-ruby (1.0.0) sha256=1a10bf312003df56ad4096e16370a4bb65b57b63366400c4404d84becea724ba
  lamby (7.0.0) sha256=d494f52b24eded8201d2c95c9d0b88cd61c5c945eec07d978e9bf94bfd4bc30e

The cost of that fix is a full re-resolution of your dependency graph on a day you did not plan to have one, which is why gem "rails", "8.1.3.1" is pinned exactly in the Gemfile used here rather than left as ~> 8.1.3. BUNDLE_FROZEN=true stays in the Dockerfile for the usual reason, that a deployed image must not silently resolve anything, but it is not what makes the boot work.

One smaller thing the read-only mount does, visible in the log of every successful boot: HOME is /root and is not writable, so bundler relocates it. `/root` is not writable. Bundler will use `/tmp/bundler20260927-17-yi4twl17' as your home directory temporarily. That line is noise here. It is not noise if something in your boot path writes to ~.

The error text you get is not your error

Every init failure in this experiment arrived wearing the same wrong face:

Init error when loading handler 
/var/runtime/gems/aws_lambda_ric-3.2.0/lib/aws_lambda_ric/lambda_logger.rb:7:in 'LambdaLogger.log_error': undefined method 'pretty_unparse' for module JSON (NoMethodError)

      puts JSON.pretty_unparse(exception.to_lambda_response)
               ^^^^^^^^^^^^^^^
Did you mean?  pretty_generate

Amazon's runtime interface client reports an initialization error by pretty-printing it, and it pretty-prints it with a method the json gem removed. json-3.0.2/CHANGES.md lists Removed JSON.pretty_unparse alongside Removed JSON.unparse and Removed Kernel#j. The base image itself is fine, because the runtime ships json 2.18.0:

$ docker run --rm --entrypoint /bin/sh public.ecr.aws/lambda/ruby:4.0 -c 'ruby -rjson -e "puts JSON::VERSION; puts JSON.respond_to?(:pretty_unparse)"'
2.18.0
true

Your bundle is what breaks it. Rails 8.1.3.1 resolves json to 3.0.2, bundler/setup puts that on the load path ahead of the default gem, and the runtime client loses a method it was written against:

$ docker run --rm --entrypoint /bin/sh lamdemo:v4 -c 'cd /var/task && BUNDLE_GEMFILE=/var/task/Gemfile ruby -e "require \"bundler/setup\"; require \"json\"; puts JSON::VERSION; puts JSON.respond_to?(:pretty_unparse)"'
3.0.2
false

The real backtrace is still printed, below the NoMethodError, which is the only reason any of this was diagnosable. Read past the first stanza in CloudWatch. The same call is present in aws_lambda_ric-3.2.0 on the ruby:3.4 base image, where the runtime's own json is 2.9.1, so the version of Ruby you pick does not change it.

Init duration, and the ten seconds you are given for it

Lambda gives CPU in proportion to memory, and the quotas page states the anchor: "At 1,769 MB, a function has the equivalent of one vCPU." So the useful axis for a Rails boot is vCPU, and docker run --cpus= is a fair proxy for it. Two runs of a fresh container at each setting, taking INIT REPORT(durationMs: ...) from the emulator's log:

--cpus roughly init run 1 init run 2
0.1 177 MB 14673 ms
0.125 221 MB 8657 ms
0.25 442 MB 4565 ms 4168 ms
0.5 885 MB 1990 ms 2057 ms
1 1769 MB 932 ms 960 ms
2 3538 MB 917 ms 889 ms
4 7076 MB 1105 ms 989 ms

Two readings. Rails boots on one core, so everything above 1769 MB is paid for and not used; the 0.89 s at 2 vCPU and the 1.1 s at 4 are the same number with noise on it. And the bottom of the table is where the real constraint lives. The execution environment lifecycle page says the Init phase "is limited to 10 seconds", and that "if all three tasks do not complete within 10 seconds, Lambda retries the Init phase at the time of the first function invocation with the configured function timeout". This function is a generated Rails app and nothing else, Bundle complete! 21 Gemfile dependencies, 87 gems now installed, and it took 14.6 seconds at 0.1 vCPU. Lambda's memory floor is 128 MB. There is no version of a Rails app that starts there.

The emulator does not enforce the 10-second cap, which is worth knowing before you trust a local run: it happily reported 14673 ms and served the request. The cap is the platform's.

For scale, this repository's own Rails application, which is a real one, boots in bin/rails runner in 1.18 s to 1.58 s on the bare metal of the same laptop with bootsnap warm. Lambda's cheapest configurations are between four and twelve times slower than that per boot, and you pay for every one of them. The same curve on an ordinary container, where the cold start is a Rails boot under a Docker CPU cap rather than an Init phase, bends the same way and at the same places, because it is the same boot.

One invocation at a time, so the thread pool is decoration

An execution environment handles a single invocation and then freezes. Set RAILS_MAX_THREADS=5, call an action that sleeps for a second, and ask the process what it thinks it is running:

{"statusCode":200, ... "body":"start=15:57:51.597 end=15:57:52.601 threads=2\n"}

Two threads: the one serving the request and the runtime client's. No Puma, no pool, no queue. The RAILS_MAX_THREADS you tuned for a dyno sets max_connections in config/database.yml and nothing else.

The emulator makes the point rudely. Two overlapping invocations against one container produce this, and then the emulator dies:

27 Sep 2026 15:57:52,835 [INFO] (rapid) ReserveFailed: AlreadyReserved
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x40 pc=0x334650]
    golang.a2z.com/LambdaRuntimeInterfaceEmulator/internal/lambda/rapidcore/server.go:666 +0xa0

That crash is the emulator's, not the platform's, and it is a reason not to load-test against it. The behaviour underneath is the same in both: real Lambda answers a concurrent request by starting another execution environment, which means another Rails boot, which means another cold start. Concurrency on Lambda is horizontal and it is paid for in init durations.

Every execution environment opens its own database connection

The first Post.count in a fresh container took 26.9 ms on one run and 58.9 ms on another; the three queries after it took 0.8 ms, 0.7 ms and 0.8 ms. That gap is the connection, and it is paid once per execution environment, exactly as the lifecycle page describes: "if your Lambda function establishes a database connection, instead of reestablishing the connection, the original connection is used in subsequent invocations". Do not read the absolute first number: the container reaches the host's PostgreSQL through a userspace TCP forwarder written for this experiment, which inflates connection setup and does not touch the queries after it. The shape is what transfers.

Which is fine until you multiply it. The default concurrency quota is 1,000. PostgreSQL here answers:

$ psql -h localhost -p 15432 -U mehdifarsi -d postgres -tAc "select name, setting from pg_settings where name in ('max_connections','superuser_reserved_connections')"
max_connections|100
superuser_reserved_connections|3

A traffic spike that opens 300 execution environments opens 300 connections to a database that accepts 97. Nothing in Rails prevents this and nothing in Lambda warns about it. The answers are a reserved concurrency limit on the function, a connection proxy in front of the database, or a database that does not care. This page does not test any of them, because none of them can be run on a laptop.

The response shape is not a Rack response

Payload format 2.0 does not have a Set-Cookie header. Two cookies set by a controller come back as a separate top-level key:

{"statusCode":200, ... "body":"two cookies\n","cookies":["a=1; path=/; samesite=lax","b=2; path=/; samesite=lax"]}

lamby-7.0.0/lib/lamby/rack_http.rb is where that split happens, and it does the format 1.0 case as multiValueHeaders in the same method. Worth knowing because a session cookie that silently does not arrive is a bug you will chase in the wrong file.

One thing rack_http.rb does that you cannot configure: ::Rack::RACK_URL_SCHEME => 'https' is hardcoded in env_base. request.ssl? is true in every invocation regardless of what the event said.

What the runtime API costs per request

Same image, same machine, same action, concurrency 1, 300 requests. Through the Lambda runtime interface, ab posting the event to the emulator:

Requests per second:    358.30 [#/sec] (mean)
Time per request:       2.791 [ms] (mean)

And the same application under Puma inside the same image, ab on /hello:

Requests per second:    529.85 [#/sec] (mean)
Time per request:       1.887 [ms] (mean)

About 0.9 ms per request to turn an event into a Rack env and the response back into JSON. That is the cost of the adapter and the local runtime API round trip. It is not the cost of API Gateway, which is not on this machine and is not measured here.

The money

Billed Duration for one cold invocation and the five warm ones after it, at --cpus=1:

REPORT RequestId: 2a52fe6d-36a6-406f-bea0-b2cafc2cc553  Init Duration: 0.16 ms  Duration: 1020.57 ms    Billed Duration: 1021 ms
REPORT RequestId: 38dbd2f3-6c1f-4178-b960-e5411eb38a33  Duration: 2.11 ms   Billed Duration: 3 ms
REPORT RequestId: c2f04423-5734-4084-a92a-ca263121f71f  Duration: 2.01 ms   Billed Duration: 3 ms
REPORT RequestId: 527b61ef-cc80-4ed3-9b5f-130f9b24e162  Duration: 2.15 ms   Billed Duration: 3 ms
REPORT RequestId: b2ca40c4-c662-4785-86a0-02b8e166a9fd  Duration: 2.09 ms   Billed Duration: 3 ms
REPORT RequestId: 5bb25aba-2a6f-4752-b362-a07e6875fbc1  Duration: 2.80 ms   Billed Duration: 3 ms

The AWS Lambda pricing page gives "$0.20 per one million requests" and, in its worked example, "$0.0000166667 per GB-s". It does not print a separate arm64 rate anywhere a reader can point at, so the arithmetic below uses the rate it does print. Run it yourself:

GB_SECOND = 0.0000166667      # AWS Lambda pricing page, x86 example rate
PER_REQUEST = 0.20 / 1_000_000
MEM_MB = 1769                 # one vCPU per the Lambda quotas page
gb = MEM_MB / 1024.0
memory 1769 MB = 1.7275 GB
one cold invocation (1021 ms billed): $0.00002960
one warm invocation (3 ms billed):   $0.00000029
ratio: 103.3x

1 req/s for 30 days = 2592000 invocations: duration $0.22 + requests $0.52 = $0.74
10 req/s for 30 days = 25920000 invocations: duration $2.24 + requests $5.18 = $7.42
100 req/s for 30 days = 259200000 invocations: duration $22.39 + requests $51.84 = $74.23

one execution environment held warm for 30 days (2_592_000 s): $74.63

Two things fall out of that block. At a request cost of $0.20 per million and a 3 ms action, more than two thirds of the bill is the per-request charge, not the compute: tuning your action from 3 ms to 2 ms changes almost nothing. And the 3 ms action is a fiction. A Rails page that touches the database is 30 ms to 100 ms, and at 50 ms the same 10 req/s costs $42.50 for the month instead of $7.42, which is where the comparison with a $7 Basic dyno stops being interesting.

The package is bigger than the zip limit

147 MB unzipped, 67,766,546 bytes zipped, for a generated Rails app with pg, nokogiri, thruster and rbs in it. The quotas page allows "50 MB (zipped, when uploaded through the Lambda API or SDKs)" and "250 MB ... including layers and custom runtimes. (unzipped)". So the unzipped side has room and the zipped side does not: a zip deploy has to be uploaded to S3 and referenced, not pushed through update-function-code. The four largest directories are the ones you would guess:

14  /var/task/vendor/bundle/ruby/4.0.0/gems/pg-1.6.3-aarch64-linux
13  /var/task/vendor/bundle/ruby/4.0.0/gems/thruster-0.1.26-aarch64-linux
11  /var/task/vendor/bundle/ruby/4.0.0/gems/nokogiri-1.19.4-aarch64-linux-gnu
10  /var/task/vendor/bundle/ruby/4.0.0/gems/rbs-4.2.0

thruster is 13 MB of Go binary that a Lambda function will never execute, and it is in the default Gemfile. rbs is 10 MB of type signatures that arrive through rdoc (8.0.0) declaring rbs (>= 4.0.0), four levels from anything you asked for. The container image was 1.2 GB against a documented ceiling of "10 GB (maximum uncompressed image size, including all layers)", so the container route removes the size question entirely, and that is the main reason to take it.

The verdict

Put a Rails app on Lambda when the traffic is spiky and mostly absent, when a two-second first response is acceptable, and when the thing on the other end is a queue consumer, a webhook receiver or an internal tool rather than a page a customer waits for. It is genuinely good at that, and the per-request arithmetic above is unbeatable at 1 req/s.

Do not put a customer-facing Rails application on it. Not because it fails, it clearly does not, but because the two properties that make Lambda cheap are the two a web application does not have: bursty traffic and idle time. At any sustained rate the bill crosses a VPS you rent by the month, and the cold start is not a tail latency problem you can tune away. It is a full Rails boot, it is on the critical path of the first request into every new execution environment, and horizontal scaling creates them by design.

What would change this: SnapStart for Ruby. The lifecycle page documents it as restoring an execution environment from a snapshot of memory and disk taken at publish time instead of running Init, which is precisely the operation that would make a Rails boot free. The runtime table lists the Ruby runtimes as ruby4.0, ruby3.4 and ruby3.3; whether SnapStart covers them is not something this page verified, and the day it does, the whole second half of this argument needs rewriting.

What this page does not cover

No function was deployed to an AWS account. Every number is from the runtime interface emulator on an M2 Max under Docker CPU limits, which is a fair way to compare configurations against each other and a bad way to predict Graviton wall-clock time. Nothing here measures API Gateway, its latency or its own pricing. Nothing here tests VPC-attached functions, where the network setup adds to cold start. Provisioned concurrency, SnapStart, Lambda Managed Instances, RDS Proxy and Active Storage on S3 are all named and none of them are tested. The 6 MB synchronous response cap is documented on the quotas page and was not exercised, so if you render a large CSV inline, find out before you ship. Solid Queue and Action Cable are not discussed at all, and neither has a sensible story on a runtime that freezes the process between invocations.

#rails #deployment #infrastructure

Comments

No comments yet. Be the first.

Only used to confirm and publish your comment. Never shown publicly, never shared.

Markdown: **bold**, `code`, ```fenced blocks```, > quotes, [links](url). HTML and images are not rendered.