Ruby on Rails on AWS Lambda
A Rails app on AWS Lambda fails at config/boot.rb line 3, and the error you are shown is a
NoMethodError inside Amazon's own runtime client rather than the thing that actually went wrong.
Everything else about this works better than it has any right to: nine lines of handler, a
Dockerfile, and a real HTTP response out of a real execution environment. The order matters, so
this page goes in the order you will meet it.
Conditions for every measurement below. Apple M2 Max, 12 cores, macOS arm64-darwin25, Docker Desktop
4.37.2, engine 27.4.0. The function is a generated Rails 8.1.3.1 app on pg, built on
public.ecr.aws/lambda/ruby:4.0 (Ruby 4.0.7, Amazon Linux 2023, aws_lambda_ric-3.2.0, bundler
4.0.20, json 2.18.0) and run through the AWS Lambda Runtime Interface Emulator that ships inside that
image at /usr/local/bin/aws-lambda-rie. PostgreSQL 17.7 on the host. Nothing here was deployed to
an AWS account: the numbers are the emulator's own REPORT lines, on this laptop's CPU, and they
are not Graviton numbers. Where a figure comes from AWS documentation rather than from a run, the
sentence says so and names the page.
The whole handler is nine lines
app.rb sits at the root of the application, which is what Lambda calls /var/task, and it is the
entire integration:
require_relative "config/boot"
require "lamby"
require_relative "config/application"
Rails.application.initialize!
def handler(event:, context:)
Lamby.handler Rails.application, event, context
end
Rails.application.initialize! runs during the Init phase, outside the handler method, which is the
one design decision that matters: boot happens once per execution environment and the invocations
after it are free of it. lamby 7.0.0 is doing the work in the last line, converting an API Gateway
event into a Rack env and the Rack triplet back into the JSON shape Lambda expects.
The Dockerfile is the other half. public.ecr.aws/lambda/ruby:4.0 already contains the runtime
interface client and its bootstrap, so there is nothing to install for Lambda itself; gcc, make
and libpq-devel are there for pg:
FROM public.ecr.aws/lambda/ruby:4.0
RUN dnf install -y gcc make libpq-devel && dnf clean all
ENV BUNDLE_GEMFILE=/var/task/Gemfile \
BUNDLE_PATH=/var/task/vendor/bundle \
BUNDLE_WITHOUT=development:test \
BUNDLE_FROZEN=true \
RAILS_ENV=production
WORKDIR /var/task
COPY Gemfile Gemfile.lock ./
RUN bundle install && bundle clean --force
COPY . .
RUN SECRET_KEY_BASE_DUMMY=1 bundle exec rails assets:precompile
CMD [ "app.handler" ]
BUNDLE_PATH=/var/task/vendor/bundle is not arbitrary. Line 10 of the image's
/var/runtime/bootstrap reads:
export GEM_PATH=${GEM_PATH:+$GEM_PATH:}/var/task/vendor/bundle/ruby/4.0.0:/opt/ruby/gems/4.0.0:/var/runtime:/var/runtime/ruby/4.0.0:/var/lang/lib/ruby/gems/4.0.0
That path is first, ahead of the runtime's own gems, which is Lambda telling you where a Ruby function's bundle is meant to live.
Run it, POST an API Gateway payload 2.0 event at the emulator, and a Rails response comes back:
$ curl -s -XPOST "http://localhost:9015/2015-03-31/functions/function/invocations" -d @event-hello.json
{"statusCode":200,"headers":{"x-frame-options":"SAMEORIGIN","x-xss-protection":"0","x-content-type-options":"nosniff","x-permitted-cross-domain-policies":"none","referrer-policy":"strict-origin-when-cross-origin","content-type":"text/plain; charset=utf-8","vary":"Accept","etag":"W/\"2f57017a45244136937a32a4f80cef35\"","cache-control":"max-age=0, private, must-revalidate","x-request-id":"699b747d-ee91-4137-a667-64bce8495f4e","x-runtime":"0.008236"},"body":"hello from pid 16 boot 1790524514.763297\n"}
One missing checksum in the lockfile stops the boot
Lambda mounts the function's code read-only and gives you /tmp and nothing else. Reproduce that
with docker run --read-only --tmpfs /tmp:rw,size=512m, and on the first attempt here the container
never reached Rails:
/var/lang/lib/ruby/4.0.0/bundler/shared_helpers.rb:122:in 'Bundler::SharedHelpers#filesystem_access': There was an error while trying to write to `/var/task/Gemfile.lock`. File system is read-only. (Bundler::ReadOnlyFileSystemError)
from /var/lang/lib/ruby/4.0.0/bundler/definition.rb:417:in 'Bundler::Definition#write_lock'
from /var/lang/lib/ruby/4.0.0/bundler/definition.rb:398:in 'Bundler::Definition#lock'
from /var/lang/lib/ruby/4.0.0/bundler/runtime.rb:104:in 'Bundler::Runtime#lock'
from /var/lang/lib/ruby/4.0.0/bundler/runtime.rb:35:in 'Bundler::Runtime#setup'
from /var/lang/lib/ruby/4.0.0/bundler.rb:166:in 'Bundler.setup'
from /var/lang/lib/ruby/4.0.0/bundler/setup.rb:32:in 'block in <top (required)>'
from /var/task/config/boot.rb:3:in 'Kernel#require'
from /var/task/config/boot.rb:3:in '<top (required)>'
Nothing in that trace is Rails. bundler/setup, on line 3 of config/boot.rb, decided the lockfile
needed rewriting. The decision is one line of runtime.rb:
def lock(opts = {})
return if @definition.no_resolve_needed?
@definition.lock(opts[:preserve_unknown_sections])
end
With a lockfile bundler is happy with, no_resolve_needed? is true, write_lock is never reached,
and the same image boots read-only with no special configuration. The lockfile in that first build
was not one bundler was happy with. It had been produced by a bundle install that added lamby to
an already-resolved lock, and that path writes the gem into CHECKSUMS without its sha256. One
entry with no checksum is a resolve, a resolve is a write_lock, and write_lock reaches this:
if File.exist?(file) && lockfiles_equal?(@lockfile_contents, contents, preserve_unknown_sections)
return if Bundler.frozen_bundle?
SharedHelpers.filesystem_access(file) { FileUtils.touch(file) }
return
end
FileUtils.touch on a read-only mount is Errno::EROFS. You can recreate the whole failure from a
working image with one substitution, which is how the chain above was confirmed rather than guessed:
RUN sed -i 's|^ lamby (7.0.0) sha256=.*| lamby (7.0.0)|' /var/task/Gemfile.lock
The dead end. return if Bundler.frozen_bundle? sits right there on the line above the touch, so
BUNDLE_FROZEN=true looks like the fix, and it is not. Frozen mode refuses the same lockfile earlier
and louder:
/var/lang/lib/ruby/4.0.0/bundler/definition.rb:491:in 'Bundler::Definition#ensure_equivalent_gemfile_and_lockfile': Your lockfile has an empty CHECKSUMS entry for "lamby", but can't be updated because frozen mode is set (Bundler::ProductionError)
bundle lock --add-checksums did not fill the gap either, and the checksums were never actually
missing from anywhere but the lockfile: curl https://index.rubygems.org/info/lamby returns
checksum:d494f52b... for 7.0.0. What worked was deleting Gemfile.lock and resolving from nothing,
in a container on the same base image so the platform-specific gems resolve for aarch64-linux:
lambda-console-ruby (1.0.0) sha256=1a10bf312003df56ad4096e16370a4bb65b57b63366400c4404d84becea724ba
lamby (7.0.0) sha256=d494f52b24eded8201d2c95c9d0b88cd61c5c945eec07d978e9bf94bfd4bc30e
The cost of that fix is a full re-resolution of your dependency graph on a day you did not plan to
have one, which is why gem "rails", "8.1.3.1" is pinned exactly in the Gemfile used here rather
than left as ~> 8.1.3. BUNDLE_FROZEN=true stays in the Dockerfile for the usual reason, that a
deployed image must not silently resolve anything, but it is not what makes the boot work.
One smaller thing the read-only mount does, visible in the log of every successful boot: HOME is
/root and is not writable, so bundler relocates it. `/root` is not writable. Bundler will use
`/tmp/bundler20260927-17-yi4twl17' as your home directory temporarily. That line is noise here. It
is not noise if something in your boot path writes to ~.
The error text you get is not your error
Every init failure in this experiment arrived wearing the same wrong face:
Init error when loading handler
/var/runtime/gems/aws_lambda_ric-3.2.0/lib/aws_lambda_ric/lambda_logger.rb:7:in 'LambdaLogger.log_error': undefined method 'pretty_unparse' for module JSON (NoMethodError)
puts JSON.pretty_unparse(exception.to_lambda_response)
^^^^^^^^^^^^^^^
Did you mean? pretty_generate
Amazon's runtime interface client reports an initialization error by pretty-printing it, and it
pretty-prints it with a method the json gem removed. json-3.0.2/CHANGES.md lists Removed
JSON.pretty_unparse alongside Removed JSON.unparse and Removed Kernel#j. The base image itself
is fine, because the runtime ships json 2.18.0:
$ docker run --rm --entrypoint /bin/sh public.ecr.aws/lambda/ruby:4.0 -c 'ruby -rjson -e "puts JSON::VERSION; puts JSON.respond_to?(:pretty_unparse)"'
2.18.0
true
Your bundle is what breaks it. Rails 8.1.3.1 resolves json to 3.0.2, bundler/setup puts that on the
load path ahead of the default gem, and the runtime client loses a method it was written against:
$ docker run --rm --entrypoint /bin/sh lamdemo:v4 -c 'cd /var/task && BUNDLE_GEMFILE=/var/task/Gemfile ruby -e "require \"bundler/setup\"; require \"json\"; puts JSON::VERSION; puts JSON.respond_to?(:pretty_unparse)"'
3.0.2
false
The real backtrace is still printed, below the NoMethodError, which is the only reason any of this
was diagnosable. Read past the first stanza in CloudWatch. The same call is present in
aws_lambda_ric-3.2.0 on the ruby:3.4 base image, where the runtime's own json is 2.9.1, so the
version of Ruby you pick does not change it.
Init duration, and the ten seconds you are given for it
Lambda gives CPU in proportion to memory, and the quotas page states the anchor: "At 1,769 MB, a
function has the equivalent of one vCPU." So the useful axis for a Rails boot is vCPU, and
docker run --cpus= is a fair proxy for it. Two runs of a fresh container at each setting, taking
INIT REPORT(durationMs: ...) from the emulator's log:
--cpus |
roughly | init run 1 | init run 2 |
|---|---|---|---|
| 0.1 | 177 MB | 14673 ms | |
| 0.125 | 221 MB | 8657 ms | |
| 0.25 | 442 MB | 4565 ms | 4168 ms |
| 0.5 | 885 MB | 1990 ms | 2057 ms |
| 1 | 1769 MB | 932 ms | 960 ms |
| 2 | 3538 MB | 917 ms | 889 ms |
| 4 | 7076 MB | 1105 ms | 989 ms |
Two readings. Rails boots on one core, so everything above 1769 MB is paid for and not used; the
0.89 s at 2 vCPU and the 1.1 s at 4 are the same number with noise on it. And the bottom of the
table is where the real constraint lives. The execution environment lifecycle page says the Init
phase "is limited to 10 seconds", and that "if all three tasks do not complete within 10 seconds,
Lambda retries the Init phase at the time of the first function invocation with the configured
function timeout". This function is a generated Rails app and nothing else, Bundle complete! 21
Gemfile dependencies, 87 gems now installed, and it took 14.6 seconds at 0.1 vCPU. Lambda's memory
floor is 128 MB. There is no version of a Rails app that starts there.
The emulator does not enforce the 10-second cap, which is worth knowing before you trust a local run: it happily reported 14673 ms and served the request. The cap is the platform's.
For scale, this repository's own Rails application, which is a real one, boots in bin/rails runner
in 1.18 s to 1.58 s on the bare metal of the same laptop with bootsnap warm. Lambda's cheapest
configurations are between four and twelve times slower than that per boot, and you pay for every
one of them. The same curve on an ordinary container, where the cold start is
a Rails boot under a Docker CPU cap rather than an Init phase, bends
the same way and at the same places, because it is the same boot.
One invocation at a time, so the thread pool is decoration
An execution environment handles a single invocation and then freezes. Set RAILS_MAX_THREADS=5,
call an action that sleeps for a second, and ask the process what it thinks it is running:
{"statusCode":200, ... "body":"start=15:57:51.597 end=15:57:52.601 threads=2\n"}
Two threads: the one serving the request and the runtime client's. No Puma, no pool, no queue. The
RAILS_MAX_THREADS you tuned for a dyno sets max_connections in config/database.yml and nothing
else.
The emulator makes the point rudely. Two overlapping invocations against one container produce this, and then the emulator dies:
27 Sep 2026 15:57:52,835 [INFO] (rapid) ReserveFailed: AlreadyReserved
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x40 pc=0x334650]
golang.a2z.com/LambdaRuntimeInterfaceEmulator/internal/lambda/rapidcore/server.go:666 +0xa0
That crash is the emulator's, not the platform's, and it is a reason not to load-test against it. The behaviour underneath is the same in both: real Lambda answers a concurrent request by starting another execution environment, which means another Rails boot, which means another cold start. Concurrency on Lambda is horizontal and it is paid for in init durations.
Every execution environment opens its own database connection
The first Post.count in a fresh container took 26.9 ms on one run and 58.9 ms on another; the
three queries after it took 0.8 ms, 0.7 ms and 0.8 ms. That gap is the connection, and it is paid
once per execution environment, exactly as the lifecycle page describes: "if your Lambda function
establishes a database connection, instead of reestablishing the connection, the original connection
is used in subsequent invocations". Do not read the absolute first number: the container reaches the
host's PostgreSQL through a userspace TCP forwarder written for this experiment, which inflates
connection setup and does not touch the queries after it. The shape is what transfers.
Which is fine until you multiply it. The default concurrency quota is 1,000. PostgreSQL here answers:
$ psql -h localhost -p 15432 -U mehdifarsi -d postgres -tAc "select name, setting from pg_settings where name in ('max_connections','superuser_reserved_connections')"
max_connections|100
superuser_reserved_connections|3
A traffic spike that opens 300 execution environments opens 300 connections to a database that accepts 97. Nothing in Rails prevents this and nothing in Lambda warns about it. The answers are a reserved concurrency limit on the function, a connection proxy in front of the database, or a database that does not care. This page does not test any of them, because none of them can be run on a laptop.
The response shape is not a Rack response
Payload format 2.0 does not have a Set-Cookie header. Two cookies set by a controller come back as
a separate top-level key:
{"statusCode":200, ... "body":"two cookies\n","cookies":["a=1; path=/; samesite=lax","b=2; path=/; samesite=lax"]}
lamby-7.0.0/lib/lamby/rack_http.rb is where that split happens, and it does the format 1.0 case as
multiValueHeaders in the same method. Worth knowing because a session cookie that silently does not
arrive is a bug you will chase in the wrong file.
One thing rack_http.rb does that you cannot configure: ::Rack::RACK_URL_SCHEME => 'https' is
hardcoded in env_base. request.ssl? is true in every invocation regardless of what the event said.
What the runtime API costs per request
Same image, same machine, same action, concurrency 1, 300 requests. Through the Lambda runtime
interface, ab posting the event to the emulator:
Requests per second: 358.30 [#/sec] (mean)
Time per request: 2.791 [ms] (mean)
And the same application under Puma inside the same image, ab on /hello:
Requests per second: 529.85 [#/sec] (mean)
Time per request: 1.887 [ms] (mean)
About 0.9 ms per request to turn an event into a Rack env and the response back into JSON. That is the cost of the adapter and the local runtime API round trip. It is not the cost of API Gateway, which is not on this machine and is not measured here.
The money
Billed Duration for one cold invocation and the five warm ones after it, at --cpus=1:
REPORT RequestId: 2a52fe6d-36a6-406f-bea0-b2cafc2cc553 Init Duration: 0.16 ms Duration: 1020.57 ms Billed Duration: 1021 ms
REPORT RequestId: 38dbd2f3-6c1f-4178-b960-e5411eb38a33 Duration: 2.11 ms Billed Duration: 3 ms
REPORT RequestId: c2f04423-5734-4084-a92a-ca263121f71f Duration: 2.01 ms Billed Duration: 3 ms
REPORT RequestId: 527b61ef-cc80-4ed3-9b5f-130f9b24e162 Duration: 2.15 ms Billed Duration: 3 ms
REPORT RequestId: b2ca40c4-c662-4785-86a0-02b8e166a9fd Duration: 2.09 ms Billed Duration: 3 ms
REPORT RequestId: 5bb25aba-2a6f-4752-b362-a07e6875fbc1 Duration: 2.80 ms Billed Duration: 3 ms
The AWS Lambda pricing page gives "$0.20 per one million requests" and, in its worked example, "$0.0000166667 per GB-s". It does not print a separate arm64 rate anywhere a reader can point at, so the arithmetic below uses the rate it does print. Run it yourself:
GB_SECOND = 0.0000166667 # AWS Lambda pricing page, x86 example rate
PER_REQUEST = 0.20 / 1_000_000
MEM_MB = 1769 # one vCPU per the Lambda quotas page
gb = MEM_MB / 1024.0
memory 1769 MB = 1.7275 GB
one cold invocation (1021 ms billed): $0.00002960
one warm invocation (3 ms billed): $0.00000029
ratio: 103.3x
1 req/s for 30 days = 2592000 invocations: duration $0.22 + requests $0.52 = $0.74
10 req/s for 30 days = 25920000 invocations: duration $2.24 + requests $5.18 = $7.42
100 req/s for 30 days = 259200000 invocations: duration $22.39 + requests $51.84 = $74.23
one execution environment held warm for 30 days (2_592_000 s): $74.63
Two things fall out of that block. At a request cost of $0.20 per million and a 3 ms action, more than two thirds of the bill is the per-request charge, not the compute: tuning your action from 3 ms to 2 ms changes almost nothing. And the 3 ms action is a fiction. A Rails page that touches the database is 30 ms to 100 ms, and at 50 ms the same 10 req/s costs $42.50 for the month instead of $7.42, which is where the comparison with a $7 Basic dyno stops being interesting.
The package is bigger than the zip limit
147 MB unzipped, 67,766,546 bytes zipped, for a generated Rails app with pg, nokogiri,
thruster and rbs in it. The quotas page allows "50 MB (zipped, when uploaded through the Lambda
API or SDKs)" and "250 MB ... including layers and custom runtimes. (unzipped)". So the unzipped
side has room and the zipped side does not: a zip deploy has to be uploaded to S3 and referenced,
not pushed through update-function-code. The four largest directories are the ones you would guess:
14 /var/task/vendor/bundle/ruby/4.0.0/gems/pg-1.6.3-aarch64-linux
13 /var/task/vendor/bundle/ruby/4.0.0/gems/thruster-0.1.26-aarch64-linux
11 /var/task/vendor/bundle/ruby/4.0.0/gems/nokogiri-1.19.4-aarch64-linux-gnu
10 /var/task/vendor/bundle/ruby/4.0.0/gems/rbs-4.2.0
thruster is 13 MB of Go binary that a Lambda function will never execute, and it is in the default
Gemfile. rbs is 10 MB of type signatures that arrive through rdoc (8.0.0) declaring
rbs (>= 4.0.0), four levels from anything you asked for. The container image was 1.2 GB against a
documented ceiling of "10 GB (maximum uncompressed image size, including all layers)", so the
container route removes the size question entirely, and that is the main reason to take it.
The verdict
Put a Rails app on Lambda when the traffic is spiky and mostly absent, when a two-second first response is acceptable, and when the thing on the other end is a queue consumer, a webhook receiver or an internal tool rather than a page a customer waits for. It is genuinely good at that, and the per-request arithmetic above is unbeatable at 1 req/s.
Do not put a customer-facing Rails application on it. Not because it fails, it clearly does not, but because the two properties that make Lambda cheap are the two a web application does not have: bursty traffic and idle time. At any sustained rate the bill crosses a VPS you rent by the month, and the cold start is not a tail latency problem you can tune away. It is a full Rails boot, it is on the critical path of the first request into every new execution environment, and horizontal scaling creates them by design.
What would change this: SnapStart for Ruby. The lifecycle page documents it as restoring an
execution environment from a snapshot of memory and disk taken at publish time instead of running
Init, which is precisely the operation that would make a Rails boot free. The runtime table lists the
Ruby runtimes as ruby4.0, ruby3.4 and ruby3.3; whether SnapStart covers them is not something
this page verified, and the day it does, the whole second half of this argument needs rewriting.
What this page does not cover
No function was deployed to an AWS account. Every number is from the runtime interface emulator on an M2 Max under Docker CPU limits, which is a fair way to compare configurations against each other and a bad way to predict Graviton wall-clock time. Nothing here measures API Gateway, its latency or its own pricing. Nothing here tests VPC-attached functions, where the network setup adds to cold start. Provisioned concurrency, SnapStart, Lambda Managed Instances, RDS Proxy and Active Storage on S3 are all named and none of them are tested. The 6 MB synchronous response cap is documented on the quotas page and was not exercised, so if you render a large CSV inline, find out before you ship. Solid Queue and Action Cable are not discussed at all, and neither has a sensible story on a runtime that freezes the process between invocations.
Comments
No comments yet. Be the first.