Skip to content

Strategy for archiving data on Google Cloud (maybe other places). #85

Description

@rmclaren

We have an important optimization problem to solve (optimizing $$ + query time [model as $$ as people time is money]).

Google Cloud has a multi-teared approach for archiving long term data (https://cloud.google.com/storage). We need to figure out our strategy for where we archive data given that data's age. It is basically a probability problem where data of a certain age will have some probability of being accessed (gamma distribution?). Accessing the data will cost money (more if we do this stupidly). The picture will probably change depending on the type of data. So basically:

  • Determine if this is actually an important problem to solve (maybe the solution is trivial)
  • Come up with a model (a cost function)
  • Figure out what data we need to collect and maybe start asking for it or collecting it (ask HPSS guys?)
  • Update the cost function based on data?
  • Come up with a practical way we could implement this.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    fetchIssue related to the access of stored dataneeds assignedTask needs assigned to one or more team membersstorageIssue related to the low level I/O or storage of data

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions