Slow Builds Lab logo
Slow Builds Lab

Public notebook for the channel

Token Maxing vs Token Saving [Thinking Outloud]

July 29, 2026

This one is more of a thinking out loud video around token maxing versus token saving. I’ve been thinking about how companies are using AI right now, especially the pressure to use more tokens, newer models, bigger models, and more automation just because it is available. But I’m not convinced more AI usage always means better work. Sometimes the smarter path might be using the right model, staying inside constraints, running slower background jobs, building local/internal models, and still doing parts of the work yourself. This is a bit of a ramble, but the main idea is simple: Not every task needs the biggest model. Not every workflow needs to burn tokens. And maybe part of learning AI is learning when not to max it out. Timestamps: 00:00 Token maxing vs token saving 01:37 Companies pushing people to use more AI 02:42 Working within token limits 04:51 Local models and internal company brains 07:19 A personal example of accidental token maxing 09:24 Not everything needs the latest model 10:39 Smaller internal AI models across companies 13:22 Staying inside constraints at work 15:25 Token cost, energy, and waste 17:46 Still doing the work yourself 20:14 Using AI inside workflows without maxing tokens 23:37 Bringing AI back in house 25:10 The new jobs around AI infrastructure 26:29 Using token maxing to build token saving 28:13 Why the future still feels exciting

Watch on YouTube

Transcript

Token Maxing vs Token Saving [Thinking Outloud]

00:00 — Token maxing vs token saving

I’m gonna try to do this one as a complete ramble because I’ve been seeing

Obviously I’ve been seeing about token maxing

But I’ve always talked about and I had this idea that tokens as we see them today is going to completely change

and I believe

basically token maxing versus token saving and that’s kind of the idea I’m thinking about because

As of now you see these companies rolling through tokens non-stop. They’re changing their models

They’re trying to every time a new model comes out. Obviously people are jumping on it

They want to use the latest and greatest because they think they need to be producing more and indeed be producing

faster

whereas

Just because the the model is newer and more advanced doesn’t mean it’s going to help you

develop better. It’s not going to help you make better, faster decisions. Sometimes even then,

it still takes time, like it still needs to process go through it, maybe a stronger model

really is going to do more prediction, more analysis, deeper research, and bring in,

take longer because it needs, it’s going to use more information that you’re not privy to. Sure,

Sure it may give you a better idea in the end, but it’s probably a good chance that

the question you were asking or the thing you were trying to get it to do using a lower

model you probably would have got the same answer because you kind of knew what you wanted

to ask, you knew what you were getting at and it didn’t require you to go full out.

01:37 — Companies pushing people to use more AI

Now this all comes about because like you see all these companies like all the layoffs

that were happening, people saying like this is how many tokens you got and one company

I have friends that work with, their biggest thing was they were trying to make people

use AI more.

So part of their evaluation or quarterly reviews was looking at how many tokens they use and

if they weren’t using enough, they needed to up it.

They needed to use more.

And then you see bigger companies blowing through the token budget like so fast.

very crazy consumption just to say they’re doing it.

It’s almost like people losing their minds,

just making sure everything they touch

flows through the AI process.

Nothing is done manually anymore or using,

it still uses their own brain.

They still have to put the prompts,

they still have to review it,

but they’re relying on it more

and they’re expected to use it more.

02:42 — Working within token limits

Now I find it funny because where I work,

I’ve always been quite impressed with how

we all seem to work within our constraints.

We didn’t set out to say we’re gonna have constraints,

but all developers, non-developers,

there’s a few salespeople, whatever.

Sometimes they get a little access

and they go a little overboard quickly,

but they learn, they learn where they went wrong,

they learn where they used too much.

don’t make it make PFS for you.

Don’t, well, you can make PFS,

don’t make it make a PowerPoint.

Don’t go crazy with the imagery or infographs

and things like that.

But there’s, again, there’s right models for that.

There are ones that I find that are very well good for that.

The only thing that I feel that we did

a little bit different is we’ve chosen a path.

We’ve chosen the one we’re gonna use and we stick to it.

I believe that different models have different places.

So I actually moving to the thing where I find

codecs to be better than cloud code in my mind.

I find it to be more efficient, better analysis.

I find it to be quicker.

And I find it to be very well, very good.

Cloud code still is top-notch

and we use that at our work all the time.

And what I’m saying is we all seem to fall within,

we don’t seem to blow our budget.

There are times when you go deep planning mode

or deep analysis.

When I’m really trying to find a bug

or make sure that everything is under,

I didn’t miss anything.

I blew through it yesterday.

There’s a lot of tests I had to do yesterday.

There was a lot of restructuring.

and I did go through it.

But that’s not a normal case for me.

Normally I can go, I don’t hit a limit at all really,

but I do use it quite a bit.

It knows me, it knows how I work.

I have my own little memory thing set up.

04:51 — Local models and internal company brains

But anyway, back to token maxing.

So what I’m trying to understand

is I’ve been talking about it for a while

where I really believe there are gonna be local models.

We are gonna have our own internal LLM

that’s running somewhere on our own server.

It doesn’t need a crazy amount of power.

It doesn’t need a ton of access outside

of what we provided and what we gave it access to

so we can learn within us, learn our company.

It doesn’t need to know what,

we don’t need to know what Coke’s doing

or we don’t need to know what 3M or some other company

how they operate, what we’re getting at is we need to know how we operate.

We need to know our guidelines, our product, our policies, and all that’s contained within

our own business.

So why would we need to outsource that in a way?

Why do we need to build a brain that runs in some cloud that we have no control over?

It makes more sense in my mind for companies to build their FAQs, their policies, their

HR policies, coding policies, their sales pitches, all their numbers, their accounting

numbers so they know their budgets, they know all these things and it’s contained within

their own system.

So in that case, you’re not token maxing.

The only maxing comes there would be the speed, so how much RAM and CPU you’re going to need,

and then how much energy is going to be used.

But other than that, you’re not blowing through tokens.

There’s just an internal process running.

And I also talked about using jobs and things like that.

So there’s no reason why a lot of the features and a lot of Jira cleanup and documents and

key pages and whatever else and jobs running, test cases, scenarios, pitches, all that could

be running in the background, low hanging fruit on a basis.

It doesn’t need to be going out and pulling in the latest model.

It doesn’t need to be attached to your credit card, running up your subscription all the

time and you know I’m at is happening right now in one of my projects my

private project from my kid I did I jumped on to jet GPT with codecs

connected my computer and I don’t know why I still don’t know what command I

ran but apparently this seems been running for days to the point that she’s

already demoed the product I haven’t looked at it yet I haven’t logged in or

nothing but I have got hit with a couple of re-ups on my, I don’t know, it’s my API.

I guess it just went overage and it’s doing like $9, $9.

It did it three or four times already and I’m only at the ninth.

07:19 — A personal example of accidental token maxing

So I told it to pause because I need her feedback at this point but the fact is it built this

product and that’s a whole other thing.

So that’s kind of token maxing.

I have something running, I have it running non-stop, I didn’t know I did.

And I put a kibosh on that real quick.

But to me, that’s token maxing.

That’s having something run non-stop with no real angle, just the purpose of it running.

I didn’t know it was running, but I guess I got it to a point that it was sustainable,

but it cost me money.

And so you can see how multiple employees on bigger projects are going mad like that.

It’s going to eat out your budget really hard, really fast.

And is the return there?

At this point in my situation, there is no return.

It’s just speed.

And then pointless speed at the end because I haven’t had feedback.

I haven’t reviewed it.

I haven’t looked at it.

Now she demoed it.

She’s put it in her Instagram feed.

done a bunch of things with it and people seem to be liking it but I haven’t seen it.

But it’s time to stop and reevaluate, get the feedback and then restart at a more token

saving situation.

09:24 — Not everything needs the latest model

But again, I really believe the cost of tokens is going to be spread out.

I don’t think we’re all going to be using the latest models.

I think this whole, it almost feels like a scam in a way to me where just because you’re

pumping out a new model, everyone jumps on board, but the token price is so high because

it’s supply and demand.

It’s the limitations, the newness of it, and everyone’s got to be on board with it.

It doesn’t have to be that way.

Why do I need to jump to, if something’s been working great for me, why don’t we need to

jump?

like the new fable just for deep diving insecurity only so I give it very tight

reins and particular jobs so if something’s very important yes I will

give it a higher model but I’m not going to max my tokens on it I’m going to

limit it and what I’m doing like I said documentation if I’m doing just a new

feature then I’m going and there’s no deadline there’s no requirement

there’s requirements but there’s no like tight tight deadline on a speed

I don’t need to be running the latest and greatest.

And I can, I still need to use my own personal brain

to go through it.

I need to build my own memory.

I need to build my own documentation and process

that goes along with it.

10:39 — Smaller internal AI models across companies

So I keep going back to the whole thing of like,

AI is going to be part of every company.

There’s going to be a group of people that understand it,

how do you, there’s going to be infrastructure around it,

security around it, different levels of it,

different types of it.

And this is a total rambling.

So I’m all over the place in a way,

but I just believe like you’re going to have

these smaller models, these smaller internal models.

You’re going to have some things distributed

throughout your system.

Like if you’re a region-based type software or a company,

then maybe there’s a place where each,

I’m just thinking out loud here.

So like within each region, it has its own copy of its corporate brain, let’s say, in

that area.

So each corporate brain has its own parts for the different policies, the different procedures,

the different coding models, the different languages, the different nuances, the customers,

the clients, all the different things that go along with that.

But then they all feed back to a central one that they all have the main terminology, the

main corporate goals and principles and policies that go in place.

and they all feed off each other,

and then you’re hooked into your call centers across,

and there’s FAQs being built,

and there’s knowledge bases,

and documentation for manuals.

There’s all the kinds of things,

and a lot of what I just said

doesn’t require to have a subscription to something.

It doesn’t require you to sign up to anthropic,

or sign up to OpenAI, or Gemini.

All it requires is for you to have your own models running.

It could be in the cloud.

That’s fine, but you can set up your own VPS

and you can have your own cloud running.

Your own LLM running in the cloud,

distribute it properly,

or you can have physical servers sitting like,

it probably wouldn’t be that,

but I have one out in the room behind me

that I use my open cloud on,

and I haven’t used that that much anymore,

but it saved me a ton.

I have Gemini running on it also.

Not Gemini, Gemma, I can’t remember what it’s called,

but it’s very slow for that.

So I limit that quite a bit.

Basically what I have it doing is I set up

so I can talk to my telegram through it

and have it run test, update linear if I needed to.

But now with JATGPT and the codecs through the phone,

just mind blown for that.

But even then, because it’s on my machine,

I have it connecting and making sure

that my open cloud is updated

so I can get updates from my telegram.

can run. Same with the book idea.

13:22 — Staying inside constraints at work

So what I’m getting at is I’m managing to do quite a bit

at work within the constraints that we have. We all are, all of us at work are not on the, none of

us except for maybe a handful of people are on the top description base. The rest of us are on

the base normal professional corporate account and very few people I know of blow through their

credits, they’re tokens and we’re all, none of us are running the latest and

greatest I don’t think. I think we’re all a little over the place and in some of

it but for the most part we use, don’t use the best, fastest model. We use the

most everyday use type model and allows us to be consistent amongst each other.

It allows us to stay within the token limits and we really don’t spend that

much money when I see what our bill is compared to you see like Microsoft just turning off

cloud code for their employees.

I see what someone like Shopify how much they spend on their tokens is pretty insane.

LinkedIn, Salesforce like these companies are blowing through money like crazy.

But to what end?

Why did they get out of this token maxing?

So I believe in token saving in my mind.

I mean like work within your constraints.

That comes from a base camp back in the day,

37 signals where like they said,

like we wanna work with what we have

until we’re busting at the seams.

Until our employees cannot handle the workload.

If every employee is working an extra four hours a week

on top of their current workload,

then that’s time to add another engineer or marketing

or whatever that department is.

Until then, until we maximize ourselves

and our current bandwidth,

we don’t need to increase our bandwidth.

And I believe that’s how we need to treat tokens.

15:25 — Token cost, energy, and waste

And I believe that token costs, what’s it attached to?

It’s attached to energy.

It’s attached to the limitations of the RAM, the CPU,

the energy in the end.

And I’ve talked with us a million times,

but like once they figure out the energy

and we start really banking on proper energy usage,

make AI, figure out how to make it better.

So I really believe there’s a lot of price gouging going on,

a lot of inflation for no reason.

I think like grocery scams,

the big bread heist in Canada

where they were upping the price

when there was no need of it.

Just because the demand was there, people wanted it,

and they could get away with it.

think StubHub or Ticketmaster,

like you’re increasing the price after-market tickets

for what purpose?

Just because you think the man’s there, people want it,

and you keep the price high.

‘Cause once you lower the price, it’s hard to go back up.

Gas, the same thing.

So let’s get the energy down, let’s get that figured out,

let’s get AI understanding how it can run more efficiently.

Let’s start building out local models

so we can save on tokens,

and not just burn through them

and not think about the consequences

of the cost of our actions,

whether it be environmental

or whether it be our pocketbooks

and ’cause companies have to justify this cost.

So why have your employees run free?

Why let them go crazy with their tokens?

I really, I don’t want to see a time

where you have to justify your token usage.

You should…

See, that’s where I’ve talked about that before too,

where you should be able to,

you should have to justify what you’re capable of doing

with the amount of tokens you’re allocated.

If you need to go above that,

then have a possible, like a good reason for it.

Don’t just burn through tokens

‘cause you don’t wanna open up your email and read it,

or you don’t wanna take the time

to do your own manual search

go find the document in your Google Drive or your OneDrive or wherever you guys keep

your documents.

Like take the time and do it yourself.

17:46 — Still doing the work yourself

Don’t just get it to blindly write you a Jira.

I just had a big Jira written.

It was a massive Jira and it wrote it and it wrote all the sub-tasks.

Now I’m not going, I’m going to read it myself and I’ve already gone through it and I’ve

made the adjustments myself.

And now I’m going to go through all the subtasks and I’m going to read those myself.

I’m not just going to make a bunch of notes and tell or one by one tell AI to fix it for

me because that’s just burning tokens for no reason and then I lose a little bit of

control and I want it’s just like the code.

I did run out of tokens yesterday and the code was at a point where I’m very used to

it doing the code review with me, doing the tests for me and running them and making sure

they’re all green but I ran out of tokens so I took the time I went file by

file read them then I tested them manually and then I did the suspect

test and I did all the checking myself I feel like I’m a toddler getting off my

train wheels telling you how it is I did all this myself I wrote it I read the

code I verified it I passed rule copy because I did it good that’s not the

case I guess in this case it is because it actually that’s what happened but you

don’t have to have it do everything for you you still you can there’s a video on

my eyes I think it’s before this one or after but you like you still have to

ride the bike is what it was called and you still have to do the thing you have

to know how to do the things and the only way you know how to do them is to

do them yourself.

If you just let AI do everything for you so you’re token maxing on everything

no matter how small the task change color. Like I had that happening with my private

projects where I’d say I want to like try getting it to try different color schemes.

I use Tailwind so I have my file and I can just easily change one file and see the changes.

I don’t need to tell AI let’s try five different colors and show me each one. Like I could

give it a prop like that, but what’s the point?

I can literally do that with a couple clicks

and a save by myself.

So that’s token saving, that’s token responsibility.

I’m not just being willy nilly with what costs,

physical money that I’m earning.

20:14 — Using AI inside workflows without maxing tokens

Like I’ve stopped, I’ve actually canceled

my cloud subscription and I had a bit of a hard time

doing it when I did it but now I don’t know anxiety over it all I feel very

confident and comfortable and actually very surprised and impressed with just

using my open AI one for everything I do and I don’t I don’t max my tokens I

don’t even know last time and I think it did yesterday because it was going by

itself and that’s like I said yes I did because it upped itself twice I think

but that should be stopped now and even like I use I use gronk for don’t use

gronk for that I use something on X but anyway like I do a lot of manual stuff

so I use N8n I built up my workflows I do all that I did all that manually

well JHPT helped me build it because it would have took me a lot longer and so

So there’s an entire workflow that happens with my X accounts and there’s only one small

piece that’s AI involved.

And then everything after that is manual.

I don’t have it doing the posts.

I don’t have it doing the follows or the reposts or anything like that.

That’s me.

I get an email.

It builds me an email and that’s it.

I don’t have it doing a bunch of analysis and going through.

I do and I have another one that’s going to go through that’s going through my emails

and doing the same thing for expenses.

And the only single part in there that that’s involved is deciphering the email, like just

reading it based off examples, it builds a knowledge base for itself to know, okay, this

email address, this subject, this is where I look, so it gets more efficient and faster

as it goes.

That’s the only part that’s AI but it’s a very big workflow that goes through a lot of different things. So it’s

using AI but using it in a token saving way rather than

Maxing it out with these massive prompts or multiple steps of massive prompts

I try to limit as much as possible and then on top of that I’m trying to use it where

It uses the local model as much as possible rather than so a new email comes in that it’s not familiar with and

Hasn’t seen before then it’ll go out and build

Decipher it then it’ll update the memory and then it’ll fall back to use the local model

Because I’m not worried about how long it takes it runs once a day it runs at night

I don’t care if it takes four hours to run

I don’t care if it takes eight hours to run because I only needed to run once and that’s fine by me

Same with the Twitter one. I don’t really care how long it takes.

It just has to run once a day and that’s it. Done.

So, and I can… so if the local model is slow, I don’t care.

But it doesn’t cost me any money besides the energy to run the machine.

It doesn’t cost me a bunch of tokens to call out and be efficient.

I don’t need it to be efficient. I need it to be consistent and accurate.

I don’t and I need it to run. That’s literally it.

23:37 — Bringing AI back in house

So it comes into the play of like what?

what people really need and

right now I don’t I don’t know cuz I only know what we do internally and

When I talk to a few friends at other companies and what I see online, but what I see online is more

People maxing all the time always pushing the new models. I don’t see a lot of time. I’m seeing more and more

more about running local, bringing stuff back in house.

But are they talking about just their data

to build their data modes?

Are they talking about getting away

from SaaS outsource solutions, the document repositories,

things like that, their code repositories?

Or are they talking about AI models bringing that in house,

which I hope they are, because I really

believe there’s a place for that.

Keep it to yourself.

You can still put it up on Amazon or DigitalOcean

or whatever and have it stored there.

That’s fine too.

But at the same time, it’s yours.

Keep it for yourself.

You don’t need to have OpenAI or Anthropic or Proplexity

or Gemini running it.

You can run your own model.

And that helps you save money there.

It helps you be more efficient.

It helps you build that data mode

that no other company would have.

It’s your own internal thing

and it gives you an advantage.

It learns from you, it knows you.

25:10 — The new jobs around AI infrastructure

Anyway, I don’t know if this is video.

This is really rambling,

’cause I really wanted to get it out there

where I really see the movement

that is either here or coming.

And the more and more we get to it,

yes, it’s gonna change the world.

Things are gonna be crazy.

We are gonna lose jobs,

but I really see a new movement of people,

new jobs being created in coming to life.

And I do believe the people that are paying attention to AI

are gonna be the leaders of those new positions,

the creators of those new positions.

They’re gonna build new positions.

They’re gonna have their own departments

that are gonna be like,

at least it’s gonna be a whole infrastructure part of it,

marketing, workflow, document management.

Like all this comes into play,

building out local models,

building like which ones, how to tweak them,

to configure them, how to make them grow, how to build those workflows and what systems

need to be in place.

There’s this whole another ecosystem of jobs that are going to come out of this, I believe,

but people had to be willing to do it.

People can’t be scared, but you can’t put your head in the sand.

You got to stay involved.

You got to learn.

So there is a little bit of token maxing where you need to try, get your feet wet, learn

how to do it when you hit the limits.

You know, don’t do that again.

Let’s keep that on the low.

26:29 — Using token maxing to build token saving

I know not like accounting it kills it when I go really really deep and just

dump all my documents all my statements everything in there and say how did I do this year? I need to farm my taxes and like

Monday

That kills it pretty quick

But it’s not that bad like that’s a one-off type of thing. That’s not like well once a year

But it’s not a crazy thing and now now it’s getting smarter

Now, if I give it the statements at the end of each month, it’s a lot less heavy.

It already has the process.

There’s a project that’s already set up.

The workflows are there.

You’re building from…

So you use a token maxing to build the processes that help you lead into the token saving.

Let AI help you be more efficient.

Let it figure out how you can be better and what parts can be human, what parts don’t

need AI, how to leverage local models, how to do it. Like how

where can I do cost savings in this case? What’s a job? What’s

a slow? What needs now? What needs right away? What’s urgent?

So I think there’s a big there’s a there’s a big market. There’s a

big place in helping companies figure this out. And even

individuals and that’ll be I think there’s something there not just giving

people open clouds given giving companies and individuals their own

workflows like what their n8 8s and 8s and their own like my dream viper and

all those things I think there’s something there and I’m again I’m so

excited about everything that’s happening I really believe that there’s

28:13 — Why the future still feels exciting

There’s the future that’s going to come out of all this

Is going to be so exciting and so much less

Hopefully it’s less stressful

Right now it’s exciting and stressful

So hopefully the stress

Goes away. It helps become better managed because once we start knowing stuff once we’re work all become comfortable

Things become a little easier to breathe and relax and very excited. I’m very always excited about this

and thanks for watching. Bye.