Using AI In Your Shed Business:
By Robert Oxley P.E., PhD, CEO of The Shed App
One techy founder. A marketing and sales team made up of AI agents. The promise? A shed company that outsells, outdelivers and outsmarts the competition down the road with forty traditional sales lots.
That is likely the pitch you are hearing right now – or at least some version of ‘a country of geniuses in a data center’. And it is not all hype. Want to help customers visualize a shed in their backyard? A prompt or two, a few seconds, and you’ve got a photo of their yard with the shed placed right in it. Need a website refresh? A weekend with the prompt window and a solid vision, and you’re live. Need an article for an industry publication? A couple weeks of prompts, some back-and-forth, a PhD in engineering, and… well, you get the picture.
And if I were to take a guess: you are not just hearing this pitch, you are already testing pieces of it. Maybe you have directed your browser to ChatGPT and asked it to give you some shed marketing ideas. Perhaps you downloaded and installed something they call Claude and prompted your way to a custom shed pricing application over the weekend. Or you’ve seen a competitor’s slick new website and felt that pinch of ‘wait, are we behind now?’
If any of that sounds familiar, welcome to the club! A lot of shed businesses are somewhere on this path, testing the waters and wondering what opportunities they are missing out on. Fear of missing out and falling behind is a powerful motivator, throw in some marketing genius and investors will come calling.
AI is an amazing tool and can do some amazing things. Rather than asking ‘What can AI do for my shed business?’, a better question is: ‘What risks am I exposing my business to when I use AI?’The goal of this article is to answer this question by providing a basic understanding of AI, what the potential risks are, and how best to limit the downside – while still getting the most out of AI amazing-ness.
Ready to get started?
AI is an amazing tool and can do some amazing things. Rather than asking ‘What can AI do for my shed business?’, a better question is: ‘What risks am I exposing my business to when I use AI?’
The goal of this article is to answer this question by providing a basic understanding of AI, what the potential risks are, and how best to limit the downside – while still getting the most out of AI amazing-ness.
Ready to get started?
What Is AI, Really?
To better understand how AI can be leveraged in your business, it’s worth taking some time to understand what AI actually is and how we got here.
AI (Artificial Intelligence) is a broad term for any computer software or system that attempts to mimic human thinking. There is a branch of AI called machine learning that aims to make accurate predictions on data that the system has never encountered before. One of the tools used in machine learning is optimization – or – perhaps better said – ‘mathematical techniques for getting the best answer’.
Before we set off on the adventure of finding the best answer to a problem, it would be good to establish several things: first – we want to have a good understanding of what good looks like. Second, we need some motivation to help us get through the challenging times. Third, we need a reliable way to not get stuck with an answer that looks pretty good in the moment, but isn’t as good as it could be.
Large Language Models – LLMs, which is really what most of us mean when we say ‘AI’ these days – use these basic ideas to optimize their internal systems (or train the model). What ‘good looks like’ for an LLM is something like ‘accurate prediction of the next word’ – when predictions are accurate, the model gets a high score. ‘Motivation’ for an LLM is chasing a perfect score, and to avoid getting stuck, LLMs use a lot of information and a bit of randomness. Repeat this enough times, and the responses start to sound… well… a bit human.
So what did all of this training actually produce? A system that’s remarkably good at one specific thing: given a context, guess the next word. And because people think with words, being able to predict the next word has led to some really cool applications. However, before we get carried away with the really cool applications, let’s spend some time familiarizing ourselves with some of the limitations that show up – limitations that are inherent in the very things that make LLMs so good at what they do.
What Good Looks Like – Plausible
Let’s say that you are a shed dealer reviewing your daily sales over the last ten Saturdays and you notice a pattern: six of the last ten Saturdays, you sold eight sheds, on two Saturdays you sold three sheds, and on the remaining two Saturdays you sold zero sheds. From high school statistics you recall that the most likely prediction for on the number of sales on a future Saturday, is eight sheds – because this occurred 60% of the time.
For good measure, you then decide to look more closely at the two Saturdays where you unfortunately sold zero sheds. These Saturdays are the only two Saturdays in your data set that happen to fall on a holiday weekend – and you realize that you can now make another prediction: on Saturdays that fall on a holiday weekend, you will sell zero sheds – and based on your limited data, this occurred 100% of the time. That seems like a pretty strong prediction – much stronger than the 60% related to selling eight sheds on a Saturday.
What does’strong prediction’ mean in this case? Does 100% mean that you are certain to not sell any sheds the next time a Saturday falls on a holiday weekend? Of course not. Another way to describe this is that it is the most plausible prediction.
As indicated earlier, LLMs are good at one thing – predicting the next word such that when the words are put together, the response ‘sounds like a human’. To do this, LLMs use the same basic ideas that we just discussed – they estimate which next words are most plausible, then use some randomness to select the next word to use in the response. This may come as a surprise at first, but when you think about it, it starts to make sense: just like you can use your shed sale data to make a prediction about how many sheds you will sell on a Saturday, if you have enough information that ‘sounds like a human’, and a computer to help you organize that information – you can start making pretty good predictions about which word comes after another word.
Also notice what happened in the shed sale example. First you predicted that selling eight sheds on a Saturday was the most plausible outcome. Then, as you collected more context, you were able to make an even more plausible prediction: that you are more likely to sell zero sheds on a Saturday that falls on a holiday weekend. LLMs work the same way. Every word they generate becomes additional context for predicting the next word, continually refining what is most plausible as the response unfolds.
(And yes, for the statisticians in the audience, two Saturdays provides far too little data to estimate this with any real confidence. Please accept my apologies.)
Recall that nothing about your shed sale prediction explains why holiday weekends affect sales. They simply capture a pattern in the data. LLMs work much the same way – they can’t reason from first principles or understand the world the way people do. The model identifies statistical patterns in language and uses those patterns to predict what comes next. And just like your shed sale prediction, the prediction that the LLM makes is not necessarily certain.
You may predict zero shed sales this Saturday, but a customer could still show up and buy a shed and your prediction was, fortunately, incorrect. The same thing happens with an LLM – the LLM generates a plausible response, but when you check it, you discover that the response is factually wrong. In the AI world, this is called a hallucination. From the model’s perspective, it has produced the sequence of words that best matches the statistical patterns it learned during training. Unlike you, the LLM has no clue that a customer could still show up and buy a shed.
Recalling the two Saturdays that you used to make a prediction on the ‘zero sales’, you can probably appreciate how having more Saturdays to look at would help you to make a better prediction. And indeed – this is the case for LLMs as well – sparse training data is one cause of hallucinations, but not the only one. Hallucinations also happen when:
- The model is asked to generate information that doesn’t exist
- It must combine multiple facts
- A prompt encourages invention
- The answer depends on information that changed after training
There is also another reason that they show up – another reason connected with trying to ‘sound more like a human’.
What Good Looks Like – Sampling
If you were to get the same response from an LLM every time you gave it a prompt, you would likely lose interest rather quickly in the responses, and probably describe them as ‘boring’ or ‘robotic’ and not very human-like. Humans stay interested in variety, and we know this intuitively. For example, you would probably naturally avoid putting the same shed picture in all of your ads or social media because you know that your potential customers will quickly lose interest in your ads and tune them out.
The same thing applies to LLM responses – if the LLM always used the ‘most mathematically plausible’ word choice, the response would (nearly) always be the same and they would not sound very human-like. To get around this, LLMs use some randomness. Instead of always picking the most plausible, they employ sampling – a random selection from a range of plausible next words. For example, if the most plausible next word is ‘shed’, but ‘building’ or ‘barn’ are also very plausible (but less so), then the LLM will sometimes choose ‘building’ or ‘barn’ instead of ‘shed’.
This sampling helps the LLM to ‘sound more creative like a human’. But it also introduces some potential problems. Using the example above, let’s say the word ‘barn’ is not quite reasonable for the next word choice – the word choice after ‘barn’ will be not quite reasonable as well, but now we are a little further off track.
The ‘photo copy of a photocopy’ analogy works well here – make a ‘photocopy of a photocopy’ often enough, and you will end up with little more than a smudge. Something similar can happen with LLMs when their own response is repeatedly used as the basis for additional responses – without outside intervention, prompting to redo, make better or fix will inevitably lead to something that is eventually unrecognizable in terms of context, meaning, voice and reasoning. Even if you prompt the model to ‘make no mistakes.’
Interestingly enough – even when we get to the end of the ‘way-off-track’ response – it still sounds entirely plausible. In fact, the response provides no indication that it is way off track, sounding just as ‘confident’ as the prior response. And the LLM is still humming along, doing what it was designed to do.
As you may well have guessed, this ‘more creative’ randomness is another path to hallucination for LLMs. Adjusting this randomness – or the temperature – for a model is a balance between ‘boring sameness’ and ‘getting off track too quickly’. The more creative the model is allowed to be, the more often it will hallucinate, and the less creative, the less like a human.
What Good Looks Like – Sycophancy
Part of the ‘sound like a human’ training for LLMs involves humans acting as ‘graders’ for the LLM response. When the human grader thinks that the response is good, the model will get a thumbs up, and for poor responses, a thumbs down. Humans are also used to rank responses for an LLM, informing the LLM as to which of the responses are preferred over others. This provides the model with feedback for modifying its responses to sound even more ‘like a human’.
As it turns out, people seem to prefer responses that validate their opinions, make themselves look smart, and that are generally agreeable. This tendency has provided some interesting feedback for LLMs to use – feedback that causes an LLM to modify its responses to be not so much more human-like as much as ‘responses that humans like’.
If you have used an LLM for any amount of time, you have likely noticed that any ideas or original thought that you include in your prompt seemingly tends to inspire the LLM to laude your brilliance and commend you for your knack at being unique and especially insightful. Sorry to break it to you, but the LLM says this to everyone.
This tendency is what is recognized as ‘sycophancy’ – and in not so polite circles, a ‘brown-noser’ or a ‘yes-man’. At some point in life, most of us begin to understand that everyone’s mother thinks that they are a genius and destined for greatness amongst their peers, and we face the brutal realities of life. Then we run into an LLM and have to come to grips all over again.
Silliness aside, the combination of plausibility, random ‘creativity’ and sycophancy can make for an especially subtle – and some would say even dangerous appeal – to our natural inclinations. There are reasons that every LLM interface includes some form of: ‘AI can make mistakes. Please double check your answers.’
None of the characteristics – Plausibility, Sampling and Sycophancy – are bugs. They are the very mechanisms that make LLMs useful. The challenge for businesses is to understand when they are acceptable and when they are not, and what safeguards can be put in place to reduce the dangers.
Risks and Safeguards
Before venturing off to dominate the market with our new found LLM insider knowledge, the question is not so much ‘What can AI do for my business?’ as much as it is ‘If AI gets this wrong, what are the consequences?’
One way to look at this is to review potential business related activities and think of the risk in terms of potential financial, reputational and legal consequences. In general – the greater the consequences of an incorrect response, the less the application should rely solely on the LLM and the more it should rely on deterministic systems and human oversight.
| Risk | Intervention Level | Activities |
|---|---|---|
| Low Consequence | Low |
|
| Moderate Consequence | Moderate |
|
| High Consequence | High |
|
(Note: this table is just an example distribution – for best results, it is highly recommended that you come up with your own distribution for your business.)
When it comes to LLM use and limiting risk exposure, safeguards can be grouped into 3 broad categories:
- Model Controls (low intervention)
- These safeguards influence the model toward better responses but do not guarantee correctness.
- System Controls (moderate intervention)
- These safeguards assume the model will sometimes be wrong and use deterministic software or external systems to validate, filter, constrain, or verify the model’s inputs, outputs, or actions.
- Human/Process Controls (high intervention)
- These safeguards remove the LLM from making high-impact decisions, relying instead on deterministic business logic or human approval before critical actions are taken.
| Safeguard | Intervention Level | Examples |
|---|---|---|
| Model Controls | Low |
|
| System Controls | Moderate |
|
| Human/Process Controls | High |
|
A Special Note on Human Safeguards
Humans-in-the-loop (HITL) deserves a special note – a human sounding tool that consistently presents a plausible, sampled and sycophantic response is a particularly seductive combination. Ever sent an AI-drafted email, received a baffling reply, and re-read your own message in horror thinking, ‘Wait… I sent WHAT?!’ Good. Neither have I.
Why does this happen? We’ve spent a lifetime relying on a simple shortcut: something that sounded plausible was also reasonable. LLMs break this shortcut.
When reviewing AI content, ‘sounds plausible’ can no longer be a proxy for ‘is reasonable’ and catching these hallucinations is likely trickier than you anticipate. Another phenomenon is what is known as Gell-Mann Amnesia Effect: ‘…where AI looks like an incredible genius synthesizing the world’s knowledge right up until you ask it about the thing you know about, then it’s an idiot.’ This is why human experience and expert opinion – consultants, engineers and professional shed builders – are more valuable than ever. Beware the ‘expert’ that suggests ‘just let the AI decide’ for any high-consequence decision.
So before you go running into your next quarterly strategy meeting convinced that you and your LLM have come up with the Mona Lisa of business strategies and are finally ready to solve all of your problems and destroy the competition… take a deep breath, present it tentatively, and let your shed selling and building team give you some feedback. Not that I would know anything personally about this… but I have heard rumors.
Conclusion
AI technology – most notably LLMs – are an amazing tool, the real value of which isn’t entirely clear at this point. They are getting a lot of attention and a lot of investment. This article is not intended to predict what the future of LLMs are – things are changing rapidly, especially with respect to the safeguards and how they can improve LLM output. Nobody knows what AI will look like in three to five years, let alone twenty. The sci-fi lover in me – the one that devoured every Asimov and Clarke novel he could get his hands on while growing up – embraces the what-ifs and coolness factor that AI raises – and as I have gotten older, the philosophical questions as well. But as a heavy daily user of LLMs with high expectations, and as someone with some academic background in optimization, there is definitely some awesome-ness with LLMs, but it seems to fall a bit short of what is being sold. And unfortunately, I think the hype is doing more harm than good.
As more companies attempt to automate their marketing assets and sales copy, it is likely that something interesting is going to happen: plausible, sampled and sycophantic LLM-generated content will become the baseline – it will be everywhere. And when synthetic communication becomes cheap, human trust, accountability and real-world craftsmanship become priceless. In fact, we are already seeing this in our marketing efforts for shed companies – images and other content generated using AI see a significant drop in interaction compared to traditional assets.
The goal isn’t to turn your shed business into an automated black box. The goal is to safely leverage AI where it is especially powerful so you and your team have more time to do what AI can never do: build real trust with real people.
Robert Oxley is the CEO of The Shed App, specializing in 3D POS and Order Fulfillment software for the shed industry. Robert did his doctoral research on the topic of systems engineering and owns a shed manufacturing company in Arizona. For more information, visit theshedapp.com or email bob@theshedapp.com
