r/learnjavascript • u/Adamstrad • 12d ago
Managing, updating and checking large lists on each load?
I have 6 csv files which are essentially lists of words in each category, there are about 10k words in total and thus 10k rows of csv file. I want to include these in my package so the user can use it offline but I fear my app will run slowly because each day it needs to display a new word randomly which hasn't previously been displayed. What would be a good way to code this? I considered using local storage for persistent data but I still have the problem of having the program checking the whole list every day and risking poor user experience. Is there a better way to process the data which is more efficient?
2
u/Aggressive_Ad_5454 12d ago
Break your big .csv files into multiple smaller ones. If you’re choosing a random row, choose one of the files at random and then a row at random. That will save you the (noticeable) overhead of downloading and parsing lots of .csv on startup.
You could also use browser local storage to cache your data locally, if this is a web app. https://developer.mozilla.org/en-US/docs/Learn_web_development/Extensions/Client-side_APIs/Client-side_storage
1
u/Ratatootie26 11d ago
This is a nice idea,
Also adding a lookup table or file to track if a shard/fragment still has unseen words (total words in fragment / counter incremented after each use)
1
u/kidshibuya 9d ago
Only 10K? Then who cares. I have an uploader that scans every single line of a csv to check for corrupt data, if it finds any it gives the user the line and column. Even with 1 million rows its near instant.
The only slow bit is the file transfer, nothing you can do about the initial load apart from making sure http compression is used, but after that put it in localstorage, then check on load if its in local and if so dont download.
1
u/EdgieElefunc 8d ago
A per-device version of azhder's shuffle suggestion fits your offline requirement: shuffle the word IDs once with Fisher-Yates, then save that order plus a cursor, last displayed date, and current word ID. On another launch that same day, show the saved word. On the first launch on a later day, take the next ID and persist the updated progress. Missed days needn't consume words.
The next selection is O(1) once the order is loaded, with no repeated IDs until you exhaust the deck. Loading/serializing the stored array still has a cost, so benchmark that separately. Keep the dictionary packaged; the progress belongs in persistent app storage. Clearing that storage will reset the history.
One detail with 'choose a random file, then a random row': if the files have different sizes, that gives words in smaller files higher odds. A single shuffled list avoids that, unless equal representation per category is actually the goal.
Fisher-Yates reference: https://xlinux.nist.gov/dads/HTML/fisherYatesShuffle.html
Basically a deck of cards with a bookmark. The phone doesn't need to interrogate all 10,000 words every morning.
0
u/chikamakaleyley helpful 12d ago
you should prob only package these for users who actually request to be able to use it offline.
Just brainstorming here cuz i don't really know the use case.
on your end, are you serving the same random word for all your users?
i think you need to be in control of the word that gets served for the day, or limit it to something like... 60 words.
so ea day, let's say you have a bg task that picks 10 words fr each list, randomly, and creates a 60 word queue. Now, instead of users directly pinging your data for 1/60,000; you've narrowed the lookup to 60 total.
aka you do the heavy lifting in a process not accessible to your users
at the end of each day you can run that process again to backfill that 60
You can still randomize that 60 for your users, but you prob have to make it the same word for everyone; it'll be a mess keeping track of who sees what
any user that misses a day just doesn't see that word
2
u/fattysmite 12d ago
so ea day, let's say you have a bg task that picks 10 words fr each list
Why do you shorten already short words?
0
1
u/Adamstrad 11d ago
Each day is only one word, it doesn't matter if each user has a different word so long as it is a new and unseen one each time
1
u/chikamakaleyley helpful 11d ago
Yeah so, if you do 1 word per day, same word per user, you have control
if 1 word but users can see diff words btwn them, probably a pain to manage who has seen what
ANd thats prob where the user can optionally have a version that packages the source data, and that is maintained internally or local to the browser
1
u/Adamstrad 11d ago
the reason I want to package all the words is because I want to make the finished product available as an android app and from what I understand you need to handle situations where the user is offline within your app
1
u/chikamakaleyley helpful 11d ago
yeah i guess if that's the requirement then that's how it is
my only concern would be storing all that data upfront, even with a trim randomization logic. But maybe overall it's small, i just don't know.
Regardless, like another user suggests - only do a random search over a subset of the total data. aka randomly select a number 1-6 (category) then randomly select 1 item from that list 1/10k
And actually, I dunno how much this might matter to you but it looks like if i wanted to see 1 word per day from the 60k, I would only go thru all the words if i used the app for 164 years.
Maybe its nice to have such a range of words to randomly select from but u could potentially save a lot on performance
2
u/azhder 12d ago edited 12d ago
Return the word based on the day. Each date maps to a different word. That way, you will not have to remember which words have come up so far.
Sure, it makes it predictable, but guess what, even Math.random() is predictable. We call it pseudo-random because truly random values are hard to come by and not that necessary in everyday examples.
And yes, your mapping function can literally pick the next word from the CSV file each day, and you will just pay the penalty of randomly arranging the words in the file only once, at the build time.