We have an important optimization problem to solve (optimizing $$ + query time [model as $$ as people time is money]).
Google Cloud has a multi-teared approach for archiving long term data (https://cloud.google.com/storage). We need to figure out our strategy for where we archive data given that data's age. It is basically a probability problem where data of a certain age will have some probability of being accessed (gamma distribution?). Accessing the data will cost money (more if we do this stupidly). The picture will probably change depending on the type of data. So basically:
- Determine if this is actually an important problem to solve (maybe the solution is trivial)
- Come up with a model (a cost function)
- Figure out what data we need to collect and maybe start asking for it or collecting it (ask HPSS guys?)
- Update the cost function based on data?
- Come up with a practical way we could implement this.
We have an important optimization problem to solve (optimizing $$ + query time [model as $$ as people time is money]).
Google Cloud has a multi-teared approach for archiving long term data (https://cloud.google.com/storage). We need to figure out our strategy for where we archive data given that data's age. It is basically a probability problem where data of a certain age will have some probability of being accessed (gamma distribution?). Accessing the data will cost money (more if we do this stupidly). The picture will probably change depending on the type of data. So basically: