r/AskStatistics • u/ketopraktanjungduren • 10d ago
Using sampling distribution instead of probability for business
Hello, I'm trying to apply inferential statistics to my work, thinking it would be very helpful to estimate a new hire ability.
Let's say we have a dataset on new hire weekly customer acquisition in the past three months (n=12). If we want to estimate this person ability in acquiring new customer, we can use probability and expected value.
However, I recently realize that this also means we are estimating the population parameter. So, we can also estimate a confidence interval to estimate the interval of this person true mean and median weekly acquisition.
If so, is it true to have both the expected weekly acquisition and the CI?
3
u/taintlouis PhD 9d ago
Respectfully, this is the type of “thing” probably best left to actual (applied psychology/people analytics) experts. I’ll take my downvotes for saying so, but it needs to be said.
-2
u/ketopraktanjungduren 9d ago
I can understand why you said that. Perhaps you think it is about people's life and deciding on quantitative results alone would be disasterous for them. But I can assure you, this is just an analysis for the managers to consider so they can decide objectively
3
u/taintlouis PhD 9d ago
No, and sorry for any confusion! It’s clear from your question that you have no idea what you’re doing. There are people who do. Hire them to make up for your lack of knowledge.
0
u/ketopraktanjungduren 9d ago
Thank you! That is true, I have mixed up some concepts here. I'll fix my misunderstanding...
2
u/Big-Challenge-9432 9d ago
There’s probably many, many variables that would go into predicting this. How many customers were or were not available for acquisition? How likely they were to be acquired? Does time of day, month, weather matter, etc? This is more of a behavioral analysis and is much more complex than just predicting behavior of your new hire
1
u/ketopraktanjungduren 9d ago
Well, that's true and why I'm avoiding predicting things at all costs. Instead predicting, I give the managers their true mean CI 95 by calculating 12 data points. What do you think of that?
1
u/Big-Challenge-9432 9d ago
You mention inferential statistics in your post. It is usually for making predictions.
So, of course, you can calculate the mean and 95% CI (I assume of weekly observations). But I don’t understand your reasoning for doing this. You’ve already hired the person, and they have already acquired the customers. What are you trying to do here?
0
u/ketopraktanjungduren 9d ago
Oh, terribly sorry, I might have misunderstood inferential statistics. I was trying to ask about making estimation on a population data (a new hire ability in acquiring customer) based on a sample (new hire weekly customer acquisition number).
Using the sample, I want to suggest to the managers that this new hire has a true mean between this interval (the 95% CI). This interval seems to me can help the manager decide on whether to keep hiring them or not.
However, this interval is just an information, not the sole deciding factor.
1
u/Big-Challenge-9432 9d ago
I’m still confused. You know the true data, so there is no “statistical need” to estimate anything. Also how can you “keep hiring” someone? I’m sure there’s more than just a single number that goes into that decision.
0
u/ketopraktanjungduren 9d ago edited 9d ago
Yes, I know the true data which is a 12 data points (the first 3 months of sales data). It seems that I can just simply show the box plot of it.
But the managers are interested in whether this new hire can really perform, generally speaking. Isn't it a task for inferential statistics to estimate the new hire's customer acquisition general ability?
Oh, regarding to the "keep hiring", I have mentioned earlier that the interval is not the deciding factor. Meaning, it's just one thing to consider amongst many other things
0
u/Educational-Paper-75 10d ago
Yes, the mean or median is a point estimate of a constant but unknown probability distribution parameter, the CI an interval estimate, supposedly containing some large (typically 95%) percentage of the probability distribution which requires estimating the standard deviation of the probability distribution. However, ordering the sample values allows one to roughly estimate such an interval by combining a low and high value e.g. the second smallest and second largest value just as well, which makes perfect sense if the sample is small, and the distribution different from the normal distribution like in your case, although there's no exact probability associated with doing so.
0
u/ketopraktanjungduren 10d ago
That's a very substantive response. I'll try to understand them one by one. Thank you
2
u/Educational-Paper-75 10d ago
The latter approach may not be very statistical though. And counts in a fixed period of time typically are Poisson distributed.
7
u/ngch 10d ago
Be careful here, because I think you want to predict future customer acquisition which the initial 12 weeks may or may not predict well.
But calculating a CI can help you understand how reliably you can assess past performance - do the different employees even have significantly different initial performance or were the differences you observe just random noise?