Skip to content

1. Initialization

John Wieczorek edited this page Oct 20, 2018 · 3 revisions

Extraction process I - Initialization

  1. Summary
  2. Overview
  3. HTTP request
  4. Process initialization
    1. store_parameters
    2. initialize_extraction
    3. Launching the next step

Summary

  1. Start with a POST request to the appropriate endpoint
  2. Request parameters are stored in memcache
  3. Basic call quality checks
    1. Proper period value
    2. Extraction not already started
  4. Create new Period entity in Datastore
  5. Clear temporary entities
  6. Call get_events process

Overview

This is the first part of the process, where parameters from the request are stored in the memcache, to be accessible throughout the workflow, some cleaning operations are performed, and the general top-level information is prepared.

HTTP request

It all starts with a POST call to the init method of the admin/parser family endpoint. Currently, the URL for that endpoint is:

http://tools-usagestats.vertnet-portal.appspot.com/admin/parser/init

According to usagestats.py, a call to this URL launches a new instance of the InitExtraction class, found in the admin.parser.InitExtraction module. This module accepts the following parameters (bold means mandatory):

  • period: the period to process (i.e., the month tocalculate) in the format 'YYYYMM'. For example, period=201811 for November 2018 usage statistics.
  • force: true/false, override existing data for the given period. Defaults to False.
  • testing: true/false, use the VertNet/statReports testing repository instead of the repositories of the publishers to store reports and send issues. Defaults to False.
  • github_store: true/false, store txt versions of the reports in GitHub repositories. Defaults to False.
  • github_issue: true/false, create new issues in GitHub to notify of report availability. Defaults to False.
  • table_name: the name of the Carto table to query in order to extract usage data. Defaults to query_log_master. Note: If table_name is provided, no time constraints will be applied to the content of the resulting reports unless the table name is the same as the default. Therefore, the table with the given name must contain all and nothing but the records for the reporting period.

Process initialization

The first steps of the process are made by the post method of class InitExtraction in module InitExtraction. This post method calls two inner methods: store_parameters and initialize_extraction.

store_parameters

This method's main function is to get the parameters of the POST call and store them, along with a set of inner parameters, in the Google Cloud Datastore.

To do so, it first makes a series of calls to self.request.get. Default values for boolean parameters are false, and the name of the Carto table to query is defined in config.py, under the CDB_TABLE variable.

        # Store call parameters
        self.period = self.request.get('period', None)
        self.force = self.request.get('force').lower() == 'true'
        self.testing = self.request.get('testing').lower() == 'true'
        self.github_store = self.request.get('github_store').lower() == 'true'
        self.github_issue = self.request.get('github_issue').lower() == 'true'
        # Get default table name, CDB_TABLE, from config.py
        self.table_name = self.request.get('table_name', CDB_TABLE)

When all parameters are obtained, the method makes a call to memcache.set_multi to store in memcache the following variables:

  • period
  • force
  • testing
  • github_store
  • github_issue
  • table_name
  • searches_extracted: whether or not the process of extracting search data has finished
  • downloads_extracted: whether or not the process of extracting download data has finished
  • processed_searches: a counter to keep track of processed search events
  • processed_downloads: a counter to keep track of processed download events

All of them with the prefix usagestats_parser_ to avoid name conflicts

        # Store in memcache for further reference
        missed = memcache.set_multi({
            # Extraction variables
            "period": self.period,
            "force": self.force,
            "testing": self.testing,
            "github_store": self.github_store,
            "github_issue": self.github_issue,
            "table_name": self.table_name,
            # Process tracking variables
            "searches_extracted": False,
            "downloads_extracted": False,
            "processed_searches": 0,
            "processed_downloads": 0
        }, key_prefix="usagestats_parser_")

The set_multi method of the memcache import returns a list with all the parameters that were not saved. This API throws a warning if any of them was not saved in memcache.

        if len(missed) > 0:
            logging.warning("Some call parameters were not added " +
                            "to the memcache: {0}".format(missed))
        else:
            logging.info("All call parameters successfully added to memcache")

initialize_extraction

After storing the parameters in the memcache, this method checks for the validity of the provided parameters and creates the basic entities to manage the full extraction process.

First, it checks if the period variable is provided and if it is well-formed:

        # Check that 'period' is provided
        if not self.period:
            logging.error("Period not found on POST body. Aborting.")
            self.error(400)
            resp = {
                "status": "error",
                "message": "Period not found on POST body. " +
                           "Aborting."
            }
            self.response.write(json.dumps(resp) + "\n")
            return 1

        # Check that 'period' is valid
        if len(self.period) != 6:
            self.error(400)
            resp = {
                "status": "error",
                "message": "Malformed period. Should be YYYYMM (e.g., 201603)"
            }
            self.response.write(json.dumps(resp) + "\n")
            return 1

Then, it tries to get any existing Period entity in the Datastore referring to the same period.

        # Get existing period
        period_key = ndb.Key("Period", self.period)
        period_entity = period_key.get()

If it exists, and the force request parameter is not provided or is set to false, the process cannot continue. If the force parameter is set to true, it will delete all the entities in that period, the period itself and will continue as it if never existed.

Be very careful when using the force parameter since it can lead to data loss and repeated issue creation.

        # If existing, abort or clear and start from scratch
        if period_entity:
            if self.force is not True:
                logging.error("Period %s already exists. " % self.period +
                              "Aborting. To override, use 'force=true'.")
                resp = {
                    "status": "error",
                    "message": "Period %s already exists. " % self.period +
                               "Aborting. To override, use 'force=true'."
                }
                self.response.write(json.dumps(resp) + "\n")
                return 1
            else:
                logging.warning("Period %s already exists. " % self.period +
                                "Overriding.")
                # Delete Reports referencing period
                r = Report.query().filter(Report.reported_period == period_key)
                to_delete = r.fetch(keys_only=True)
                logging.info("Deleting %d Report entities" % len(to_delete))
                deleted = ndb.delete_multi(to_delete)
                logging.info("%d Report entities removed" % len(deleted))

                # Delete Period itself
                logging.info("Deleting Period %s" % period_key)
                period_key.delete()
                logging.info("Period entity deleted")

Finally, it creates the new period and initializes some of its variables:

  • year
  • month
  • status: will keep track of the status of the whole extraction. It is set to in progress at this point
        # Create new Period (id=YYYYMM)
        logging.info("Creating new Period %s" % self.period)
        y, m = (int(self.period[:4]), int(self.period[-2:]))
        p = Period(id=self.period)
        p.year = y
        p.month = m
        p.status = 'in progress'
        period_key = p.put()

        # Check
        if period_key:
            logging.info("New Period %s created successfully." % self.period)
            logging.info("New period's key = %s" % period_key)
        else:
            self.error(500)
            logging.error("Could not create new Period %s" % self.period)
            resp = {
                "status": "error",
                "message": "Could not create new Period %s" % self.period
            }
            self.response.write(json.dumps(resp) + "\n")
            return 1

If any previous attempt at extracting data for this period failed and any temporary information was still stored, the last part of this method cleans that up.

Reports that are being processed are stored as entities of the ReportToProcess kind. A start from scratch doesn't need any of these.

        # Clear temporary entities
        keys_to_delete = ReportToProcess.query().fetch(keys_only=True)
        logging.info("Deleting %d temporal (internal use only) entities"
                     % len(keys_to_delete))
        ndb.delete_multi(keys_to_delete)

Launching the next step

If everything went correctly, the post method finishes with a deferred call to the next method get_events. The endpoint for this method is stored in the URI_GET_EVENTS variable in the config.py module.

Also, the Usage Stats Generator has its own task queue, defined in the QUEUENAME variable in the config.py module.

        # Create task for extracting events
        taskqueue.add(url=URI_GET_EVENTS,
                      queue_name=QUEUENAME)

Finally, it creates a JSON response to the user and finishes the request.

        # Build response
        resp = {
            "status": "success",
            "message": "Period initialized and extractions enqueued",
            "data": {
                "period": self.period
            }
        }
        self.response.write(json.dumps(resp) + "\n")

Clone this wiki locally