Where ruby_llm 2.0 keeps token counts, and the dashboard that read zero
September 21, 2026
An upgrade that moves where a gem stores a number is the most dangerous kind of change, because the
old column is still there and still sums to something. In ruby_llm 2.0 that number is the token
count, and the something it sums to is zero.
Rails LLM cost reporting depends on reading the right table, and ruby_llm token usage changed tables
between major versions without changing the columns it left behind.
What moved
Before 2.0, a message carried its own token counts and a dashboard summed them off messages. In
2.0 the gem writes a row to ruby_llm_usages for every provider attempt and never sets a token
count on a message again.
Cache tokens are broken out separately because prompt caching is billed differently from fresh
input, and the cost columns carry ten decimal places because per-token prices are small enough that
rounding at the row destroys the total. operation is constrained to a list that includes
embeddings, images and transcription, so a single table covers every kind of call the gem can make
rather than chat alone.
The regression this causes, and why it is silent
A sum over a column that exists and is always null returns zero. Nothing raises. The page renders,
the tests that assert a 200 pass, and the number is simply wrong from the moment the upgrade ships.
That is exactly what happened to the admin dashboard here. It read tokens off messages, the
upgrade landed, and every token figure on the screen became zero while every other number on the
same page stayed correct. A missing method would have been better: it would have raised on the
first visit and been fixed in minutes.
The fix is to sum the new table, and it is worth keeping the reason in the code rather than in a
commit message nobody reads:
COALESCE on both sides because a row can carry one and not the other, and a null anywhere in the
addition makes the whole expression null rather than the other operand.
Joining a polymorphic table with no association
ruby_llm_usages belongs to its chat through chat_id and chat_type, and the application's
Chat model has no has_many :usages. So a query that needs chat data, such as bucketing tokens by
the month the chat was created, has no association to join through and the SQL is written by hand:
CHATS_JOIN=<<~SQL.squish.freeze
INNER JOIN chats ON chats.id = ruby_llm_usages.chat_id
AND ruby_llm_usages.chat_type = 'Chat'
SQL
The type condition is the part to keep. With one acts_as_chat model the join is correct without
it, and the day a second model acts as a chat, ids collide across the two tables and rows start
attaching to the wrong records. A condition that costs nothing today and prevents a silent data bug
later is worth the line.
The constant that breaks eager loading
Reaching the usage model has one trap worth naming, because it fails at boot rather than at runtime.
RubyLLM::ActiveRecord::Usage is defined from an on_load :active_record hook, which has not
necessarily fired when eager loading reaches an application class.
Written as a constant in a class body, the reference is resolved while the class is being loaded and
raises NameError in any environment that eager loads, which is production and the CI run and not
the development server. A method defers the lookup to the moment it is called, by which time the
hook has long since run.
Which tokens count
A usage row is written per attempt, and status distinguishes what happened to that attempt. A
request that hit a rate limit after the provider had already processed the input still produced
input tokens, and the provider charges for them.
So "how many tokens did we use" and "how many tokens produced an answer" are different questions
with different filters, and a dashboard that sums every row is answering the first. That is the
right default for a cost figure and the wrong one for a quality figure, and the difference is
invisible on the screen unless the page says which it is showing.
The failover path makes this concrete: a primary model that fails and a backup that succeeds write
two usage rows for one visible answer.
The two fallbacks covers when
that second attempt happens.
The quota a user sees is a different number
None of the above is what limits an individual. Ai::UsageQuota counts chats, not tokens, and does
it per user:
One completion is one chat, so the allowance a user sees is a count of things they did rather than a
number of tokens they cannot estimate. It is enforced in the controller before the job is enqueued,
which is early enough that an over-quota user costs nothing at all.
The trade is that a user who sends 8000-character prompts and a user who sends one line get the same
allowance, and the input cap in the controller is what stops the first one from being unbounded.
Cost control across the whole application is the model quota instead, and the two are checked in
different places for different reasons.
What this page does not cover
Turning the cost columns into a bill. The dashboard here multiplies tokens by a flat estimated rate,
which is honest about being an estimate, while ruby_llm_usages already carries a real per-row cost
that varies by model. Summing those instead is the accurate version and nothing in this codebase
does it yet.
Nor does it cover the streaming path's usage rows, which is where the tokens come from in the first
place. The job and its broadcasts are in
streaming an LLM answer with Turbo Streams.
An upgrade that moves where a gem stores a number is the most dangerous kind of change, because the old column is still there and still sums to something. In
ruby_llm2.0 that number is the token count, and the something it sums to is zero.Rails LLM cost reporting depends on reading the right table, and ruby_llm token usage changed tables between major versions without changing the columns it left behind.
What moved
Before 2.0, a message carried its own token counts and a dashboard summed them off
messages. In 2.0 the gem writes a row toruby_llm_usagesfor every provider attempt and never sets a token count on a message again.The table is wider than the columns it replaced:
Cache tokens are broken out separately because prompt caching is billed differently from fresh input, and the cost columns carry ten decimal places because per-token prices are small enough that rounding at the row destroys the total.
operationis constrained to a list that includes embeddings, images and transcription, so a single table covers every kind of call the gem can make rather than chat alone.The regression this causes, and why it is silent
A sum over a column that exists and is always null returns zero. Nothing raises. The page renders, the tests that assert a 200 pass, and the number is simply wrong from the moment the upgrade ships.
That is exactly what happened to the admin dashboard here. It read tokens off
messages, the upgrade landed, and every token figure on the screen became zero while every other number on the same page stayed correct. A missing method would have been better: it would have raised on the first visit and been fixed in minutes.The fix is to sum the new table, and it is worth keeping the reason in the code rather than in a commit message nobody reads:
COALESCEon both sides because a row can carry one and not the other, and a null anywhere in the addition makes the whole expression null rather than the other operand.Joining a polymorphic table with no association
ruby_llm_usagesbelongs to its chat throughchat_idandchat_type, and the application'sChatmodel has nohas_many :usages. So a query that needs chat data, such as bucketing tokens by the month the chat was created, has no association to join through and the SQL is written by hand:The type condition is the part to keep. With one
acts_as_chatmodel the join is correct without it, and the day a second model acts as a chat, ids collide across the two tables and rows start attaching to the wrong records. A condition that costs nothing today and prevents a silent data bug later is worth the line.The constant that breaks eager loading
Reaching the usage model has one trap worth naming, because it fails at boot rather than at runtime.
RubyLLM::ActiveRecord::Usageis defined from anon_load :active_recordhook, which has not necessarily fired when eager loading reaches an application class.Written as a constant in a class body, the reference is resolved while the class is being loaded and raises
NameErrorin any environment that eager loads, which is production and the CI run and not the development server. A method defers the lookup to the moment it is called, by which time the hook has long since run.Which tokens count
A usage row is written per attempt, and
statusdistinguishes what happened to that attempt. A request that hit a rate limit after the provider had already processed the input still produced input tokens, and the provider charges for them.So "how many tokens did we use" and "how many tokens produced an answer" are different questions with different filters, and a dashboard that sums every row is answering the first. That is the right default for a cost figure and the wrong one for a quality figure, and the difference is invisible on the screen unless the page says which it is showing.
The failover path makes this concrete: a primary model that fails and a backup that succeeds write two usage rows for one visible answer. The two fallbacks covers when that second attempt happens.
The quota a user sees is a different number
None of the above is what limits an individual.
Ai::UsageQuotacounts chats, not tokens, and does it per user:One completion is one chat, so the allowance a user sees is a count of things they did rather than a number of tokens they cannot estimate. It is enforced in the controller before the job is enqueued, which is early enough that an over-quota user costs nothing at all.
The trade is that a user who sends 8000-character prompts and a user who sends one line get the same allowance, and the input cap in the controller is what stops the first one from being unbounded. Cost control across the whole application is the model quota instead, and the two are checked in different places for different reasons.
What this page does not cover
Turning the cost columns into a bill. The dashboard here multiplies tokens by a flat estimated rate, which is honest about being an estimate, while
ruby_llm_usagesalready carries a real per-row cost that varies by model. Summing those instead is the accurate version and nothing in this codebase does it yet.Nor does it cover the streaming path's usage rows, which is where the tokens come from in the first place. The job and its broadcasts are in streaming an LLM answer with Turbo Streams.