Job request: 10933

Organisation:: University of Bristol
Workspace:: vaccine-effectiveness-in-children
ID:: jq76cd26fn3xrfkt

This page shows the technical details of what happened when the authorised researcher Will Hulme requested one or more actions to be run against real patient data within a secure environment.

By cross-referencing the list of jobs with the pipeline section below, you can infer what security level the outputs were written to.

The output security levels are:

highly_sensitive
- Researchers can never directly view these outputs
- Researchers can only request code is run against them
moderately_sensitive
- Can be viewed by an approved researcher by logging into a highly secure environment
- These are the only outputs that can be requested for public release via a controlled output review service.

Jobs

Action:

matching1

Status:

Status: Succeeded

Job identifier:

ao6gzrcslbyaddbg
Action:

skim_potential_matched1

Status:

Status: Succeeded

Job identifier:

ttdgnbjidqisqe3e
Action:

generate_study_match_control1

Status:

Status: Succeeded

Job identifier:

puyzox3b2uhggn46
Action:

matching_filter1

Status:

Status: Failed

Job identifier:

t5rrh3o4ge24tcz2

Error:

nonzero_exit: Job exited with error code 1
Action:

generate_dataset_report_control_match1

Status:

Status: Succeeded

Job identifier:

425gfiovsjkfjdck

Pipeline

Show project.yaml

version: '3.0'

expectations:
  population_size: 1000

actions:

  generate_study_treated:
    run: cohortextractor:latest generate_cohort --study-definition study_definition_treated  --output-format feather
    outputs:
      highly_sensitive:
        cohort: output/input_treated.feather

  generate_dataset_report:
    run: >
      dataset-report:v0.0.9
        --input-files output/input_treated.feather
        --output-dir output/
    needs: [generate_study_treated]
    outputs:
      moderately_sensitive:
        dataset_report: output/input_treated.html

  data_process_treated:
    run: r:latest analysis/data_process_treated.R over12
    needs:
    - generate_study_treated
    outputs:
      highly_sensitive:
        rds: output/data/data_treated_eligible.rds
      moderately_sensitive:
        flowchart: output/data/flowchart_treated_eligible.csv

# Skim data
  skim_data_treated:
    run: r:latest analysis/data_skim.R output/data/data_treated_eligible.rds output/data_properties
    needs: [data_process_treated]
    outputs: 
      moderately_sensitive:
        txt1: output/data_properties/data_treated_eligible_skim.txt
        txt2: output/data_properties/data_treated_eligible_coltypes.txt
        txt3: output/data_properties/data_treated_eligible_tabulate.txt

  ## matching round 1
  
  generate_study_control_potential1:
    run: cohortextractor:latest generate_cohort --study-definition study_definition_control_potential --output-format feather --index-date-range "2021-09-20 to 2021-09-20 by week"
    outputs:
      highly_sensitive:
        cohort: output/input_control_potential_2021-09-20.feather

 
  data_process_control_potential1:
    run: r:latest analysis/data_process_control.R over12 1
    needs:
    - generate_study_control_potential1
    outputs:
      highly_sensitive:
        rds: output/data/data_control_potential1.rds
 
  matching1:
    run: r:latest analysis/matching.R over12 1
    needs:
    - data_process_treated
    - data_process_control_potential1
    outputs:
      highly_sensitive:
        rds1: output/match/data_potential_matchstatus1.rds
        rds2: output/match/data_potential_matched1.rds
        csv: output/match/potential_matched_controls1.csv.gz
 
  skim_potential_matched1:
    run: r:latest analysis/data_skim.R output/match/data_potential_matched1.rds output/data_properties
    needs: [matching1]
    outputs: 
      moderately_sensitive:
        txt1: output/data_properties/data_potential_matched1_skim.txt
        txt2: output/data_properties/data_potential_matched1_coltypes.txt
        txt3: output/data_properties/data_potential_matched1_tabulate.txt
 
  generate_study_match_control1:
    run: cohortextractor:latest generate_cohort --study-definition study_definition_control_match1 --output-format feather
    needs:
    - matching1
    outputs:
      highly_sensitive:
        cohort: output/input_control_match1.feather 

  generate_dataset_report_control_match1:
    run: >
      dataset-report:v0.0.9
        --input-files output/input_control_match1.feather
        --output-dir output/
    needs: [generate_study_match_control1]
    outputs:
      moderately_sensitive:
        dataset_report: output/input_control_match1.html
  
  matching_filter1:
    run: r:latest analysis/matching_filter.R over12 1
    needs:
    - matching1
    - generate_study_match_control1
    outputs:
      highly_sensitive:
        rds1: output/match/data_matchstatus_allrounds1.rds
        rds2: output/match/data_match_actual1.rds
        
        
#   ## matching round 2
  
#   generate_study_control_potential1:
#     run: cohortextractor:latest generate_cohort --study-definition study_definition_control_potential --output-format feather --index-date-range "2021-10-04 to 2021-10-04 by week"
#     outputs:
#       highly_sensitive:
#         cohort: output/input_control_potential_2021-10-04.feather

 
#   data_process_control_potential2:
#     run: r:latest analysis/data_process_control.R over12 2
#     needs:
#     - generate_study_control_potential2
#     outputs:
#       highly_sensitive:
#         rds: output/data/data_control_potential2.rds
 
#   matching2:
#     run: r:latest analysis/matching.R over12 2
#     needs:
#     - data_matching_filter1
#     - data_process_treated
#     - data_process_control_potential2
#     outputs:
#       highly_sensitive:
#         rds1: output/match/data_potential_matchstatus2.rds
#         rds2: output/match/data_potential_matched2.rds
#         csv: output/match/potential_matched_controls2.csv.gz
 
 
#   generate_study_match_control2:
#     run: cohortextractor:latest generate_cohort --study-definition study_definition_control_match2 --output-format feather
#     needs:
#     - matching2
#     outputs:
#       highly_sensitive:
#         cohort: output/input_control_match2.feather 
  
#   matching_filter2:
#     run: r:latest analysis/matching_filter.R over12 2
#     needs:
#     - matching_filter1
#     - matching2
#     - generate_study_match_control2
#     outputs:
#       highly_sensitive:
#         rds1: output/match/data_matchstatus_allrounds2.rds
#         rds2: output/match/data_match_actual2.rds


# ## combine together

#   matching_combine:
#     run: r:latest analysis/matching_combine.R over12
#     needs:
#     - matching_filter1
#     - matching_filter2
#     outputs:
#       highly_sensitive:
#         rds: output/match/data_match_all.rds

Timeline

Created: 3 years, 6 months ago 18 Aug 2022 07:50:35 UTC
Started: 3 years, 6 months ago 18 Aug 2022 07:50:39 UTC
Finished: 3 years, 6 months ago 18 Aug 2022 08:48:25 UTC
Runtime: 00:58:19

These timestamps are generated and stored using the UTC timezone on the TPP backend.

Job request

Status: Failed
Backend: TPP
Workspace: vaccine-effectiveness-in-children
Requested by: Will Hulme
Branch: main
Force run dependencies: No
Git commit hash: f137782
Requested actions: matching1

skim_potential_matched1

generate_study_match_control1

generate_dataset_report_control_match1

matching_filter1

Code comparison

Compare the code used in this job request