{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.11"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":101849,"databundleVersionId":12846694,"sourceType":"competition"}],"dockerImageVersionId":31040,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false},"papermill":{"default_parameters":{},"duration":123.041003,"end_time":"2024-08-29T13:27:04.243126","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-08-29T13:25:01.202123","version":"2.5.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Calibrating and Binning Ariel Data","metadata":{"papermill":{"duration":0.013475,"end_time":"2024-08-29T13:25:05.231612","exception":false,"start_time":"2024-08-29T13:25:05.218137","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"**UPDATE**\nN/A","metadata":{"papermill":{"duration":0.012946,"end_time":"2024-08-29T13:25:05.25681","exception":false,"start_time":"2024-08-29T13:25:05.243864","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# PLEAES READ BEFORE YOU PROCEED\n1. within the wall time. From experience, GPU acceleration, Parallelisation and other coding language will help. We encourage the kaggle community to share their calibration scripts to help speed up this process.  \n\n2. We have also only processed the 1st observation of every planet, not their repeats, so you might want to also take that into account. \n\n3. While these steps help reduce noise and data size, they may not be the most effective approach for achieving the optimal model for this challenge. Participants are encouraged to explore different combinations of these steps, or other approaches, to clean up the data.","metadata":{}},{"cell_type":"markdown","source":"This notebook follows largely from the [ADC2024 edition](https://www.kaggle.com/code/gordonyip/update-calibrating-and-binning-astronomical-data), with minor updates to reflect changes made. \n\nData reduction is crucial in astronomical observations. This notebook outlines essential calibration steps typically employed by astronomers to mitigate noise in data. We want to emphasis that these steps are NOT laws or rules you must take. Take them as a general recipe where you can always add/delete things. \n\nKey points:\n\n- The notebook guides participants through pre-processing data and saving it in a more convenient, lighter format.\n- Similar to last year, if you plan to use the baseline models (which will be released soon), you must run this notebook first before training.\n\n\n","metadata":{"papermill":{"duration":0.011873,"end_time":"2024-08-29T13:25:05.280847","exception":false,"start_time":"2024-08-29T13:25:05.268974","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"\n**Acknowledgement**: This notebook is modified by Gordon Yip from its 2024 version prepared by Angèle Syty and Virginie Batista (IAP), with support from Andrea Bocchieri (Sapienza Università di Roma\n), Orphée Faucoz (CNES), Lorenzo V. Mugnai (Cardiff University & UCL), Tara Tahseen (UCL) and the Kaggle community.","metadata":{"papermill":{"duration":0.011863,"end_time":"2024-08-29T13:25:05.305762","exception":false,"start_time":"2024-08-29T13:25:05.293899","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Last modified: 27th Jun 2025. ","metadata":{"papermill":{"duration":0.011996,"end_time":"2024-08-29T13:25:05.330232","exception":false,"start_time":"2024-08-29T13:25:05.318236","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport itertools\nimport os\nimport glob \nfrom astropy.stats import sigma_clip\n\nfrom tqdm import tqdm","metadata":{"papermill":{"duration":1.619176,"end_time":"2024-08-29T13:25:06.963493","exception":false,"start_time":"2024-08-29T13:25:05.344317","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:07:58.997569Z","iopub.execute_input":"2025-06-27T08:07:58.997888Z","iopub.status.idle":"2025-06-27T08:08:01.024961Z","shell.execute_reply.started":"2025-06-27T08:07:58.997865Z","shell.execute_reply":"2025-06-27T08:08:01.023661Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Below, we define the corrections we want to apply, the size of the data chunks and the different path used to import data and save the light ones. ","metadata":{"papermill":{"duration":0.012487,"end_time":"2024-08-29T13:25:06.988825","exception":false,"start_time":"2024-08-29T13:25:06.976338","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\npath_folder = '/kaggle/input/ariel-data-challenge-2025/' # path to the folder containing the data\npath_out = '/kaggle/tmp/data_light_raw/' # path to the folder to store the light data\noutput_dir = '/kaggle/tmp/data_light_raw/' # path for the output directory","metadata":{"papermill":{"duration":0.021737,"end_time":"2024-08-29T13:25:07.023919","exception":false,"start_time":"2024-08-29T13:25:07.002182","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.028045Z","iopub.execute_input":"2025-06-27T08:08:01.029237Z","iopub.status.idle":"2025-06-27T08:08:01.036427Z","shell.execute_reply.started":"2025-06-27T08:08:01.029201Z","shell.execute_reply":"2025-06-27T08:08:01.03549Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"If the *path_out* folder doesn't exist yet, it is created. ","metadata":{"papermill":{"duration":0.012771,"end_time":"2024-08-29T13:25:07.049096","exception":false,"start_time":"2024-08-29T13:25:07.036325","status":"completed"},"tags":[]}},{"cell_type":"code","source":"if not os.path.exists(path_out):\n    os.makedirs(path_out)\n    print(f\"Directory {path_out} created.\")\nelse:\n    print(f\"Directory {path_out} already exists.\")\n","metadata":{"papermill":{"duration":0.024159,"end_time":"2024-08-29T13:25:07.08616","exception":false,"start_time":"2024-08-29T13:25:07.062001","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.037358Z","iopub.execute_input":"2025-06-27T08:08:01.037689Z","iopub.status.idle":"2025-06-27T08:08:01.098944Z","shell.execute_reply.started":"2025-06-27T08:08:01.037651Z","shell.execute_reply":"2025-06-27T08:08:01.096729Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"**Data import:**\n\n The files are imported by chunks of size 'CHUNK_SIZE' to avoid exceeding the memory capacity. ","metadata":{"papermill":{"duration":0.012256,"end_time":"2024-08-29T13:25:07.112095","exception":false,"start_time":"2024-08-29T13:25:07.099839","status":"completed"},"tags":[]}},{"cell_type":"code","source":"CHUNKS_SIZE = 1","metadata":{"papermill":{"duration":0.021663,"end_time":"2024-08-29T13:25:07.146745","exception":false,"start_time":"2024-08-29T13:25:07.125082","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.100712Z","iopub.execute_input":"2025-06-27T08:08:01.101047Z","iopub.status.idle":"2025-06-27T08:08:01.13453Z","shell.execute_reply.started":"2025-06-27T08:08:01.101012Z","shell.execute_reply":"2025-06-27T08:08:01.132855Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 1: Analog-to-Digital Conversion\n\nThe Analog-to-Digital Conversion (adc) is performed by the detector to convert the pixel voltage into an integer number. We revert this operation by using the gain and offset for the calibration files 'train_adc_info.csv'.\n","metadata":{"execution":{"iopub.execute_input":"2024-08-02T14:38:41.270597Z","iopub.status.busy":"2024-08-02T14:38:41.270117Z","iopub.status.idle":"2024-08-02T14:38:41.303633Z","shell.execute_reply":"2024-08-02T14:38:41.30235Z","shell.execute_reply.started":"2024-08-02T14:38:41.270558Z"},"papermill":{"duration":0.01211,"end_time":"2024-08-29T13:25:07.171116","exception":false,"start_time":"2024-08-29T13:25:07.159006","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def ADC_convert(signal, gain=0.4369, offset=-1000):\n    \"\"\"The Analog-to-Digital Conversion (adc) is performed by the detector to convert\n    the pixel voltage into an integer number. Since we are using the same conversion number \n    this year, we have simply hard-coded it inside. \"\"\"\n    signal = signal.astype(np.float64)\n    signal /= gain\n    signal += offset\n    return signal","metadata":{"papermill":{"duration":0.023531,"end_time":"2024-08-29T13:25:07.207817","exception":false,"start_time":"2024-08-29T13:25:07.184286","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.135958Z","iopub.execute_input":"2025-06-27T08:08:01.136321Z","iopub.status.idle":"2025-06-27T08:08:01.167134Z","shell.execute_reply.started":"2025-06-27T08:08:01.13629Z","shell.execute_reply":"2025-06-27T08:08:01.165303Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: Mask hot/dead pixel\nThe dead pixels map is a map of the pixels that do not respond to light and, thus, can’t be accounted for any calculation. In all these frames the dead pixels are masked using python masked arrays. The bad pixels are thus masked but left uncorrected. Some methods can be used to correct bad-pixels but this task, if needed, is left to the participants.","metadata":{"papermill":{"duration":0.011774,"end_time":"2024-08-29T13:25:07.232416","exception":false,"start_time":"2024-08-29T13:25:07.220642","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def mask_hot_dead(signal, dead, dark):\n    hot = sigma_clip(\n        dark, sigma=5, maxiters=5\n    ).mask\n    hot = np.tile(hot, (signal.shape[0], 1, 1))\n    dead = np.tile(dead, (signal.shape[0], 1, 1))\n    signal = np.ma.masked_where(dead, signal)\n    signal = np.ma.masked_where(hot, signal)\n    return signal","metadata":{"papermill":{"duration":0.023597,"end_time":"2024-08-29T13:25:07.268543","exception":false,"start_time":"2024-08-29T13:25:07.244946","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.16903Z","iopub.execute_input":"2025-06-27T08:08:01.169558Z","iopub.status.idle":"2025-06-27T08:08:01.203139Z","shell.execute_reply.started":"2025-06-27T08:08:01.16952Z","shell.execute_reply":"2025-06-27T08:08:01.201304Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 2: linearity Correction","metadata":{"papermill":{"duration":0.013391,"end_time":"2024-08-29T13:25:07.29542","exception":false,"start_time":"2024-08-29T13:25:07.282029","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"\n\n**Non-linearity of pixels' response:**\n\nThe non-linearity of the pixels’ response can be explained as capacitive leakage on the readout electronics of each pixel during the integration time. The number of electrons in the well is proportional to the number of photons that hit the pixel, with a quantum efficiency coefficient. However, the response of the pixel is not linear with the number of electrons in the well. This effect can be described by a polynomial function of the number of electrons actually in the well. The data is provided with calibration files linear_corr.parquet that are the coefficients of the inverse polynomial function and can be used to correct this non-linearity effect.\n\n","metadata":{"papermill":{"duration":0.014964,"end_time":"2024-08-29T13:25:07.327644","exception":false,"start_time":"2024-08-29T13:25:07.31268","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def apply_linear_corr(linear_corr,clean_signal):\n    linear_corr = np.flip(linear_corr, axis=0)\n    for x, y in itertools.product(\n                range(clean_signal.shape[1]), range(clean_signal.shape[2])\n            ):\n        poli = np.poly1d(linear_corr[:, x, y])\n        clean_signal[:, x, y] = poli(clean_signal[:, x, y])\n    return clean_signal\n    ","metadata":{"papermill":{"duration":0.032924,"end_time":"2024-08-29T13:25:07.375474","exception":false,"start_time":"2024-08-29T13:25:07.34255","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.206994Z","iopub.execute_input":"2025-06-27T08:08:01.207317Z","iopub.status.idle":"2025-06-27T08:08:01.237081Z","shell.execute_reply.started":"2025-06-27T08:08:01.207289Z","shell.execute_reply":"2025-06-27T08:08:01.233044Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 3: dark current subtraction\n\nThe data provided include calibration for dark current estimation, which can be used to pre-process the observations. Dark current represents a constant signal that accumulates in each pixel during the integration time, independent of the incoming light. To obtain the corrected image, the following conventional approach is applied: The data provided include calibration files such as dark frames or dead pixels' maps. They can be used to pre-process the observations. The dark frame is a map of the detector response to a very short exposure time, to correct for the dark current of the detector.\n$$\\text{image - dark} \\times \\Delta t $$ \nThe corrected image is conventionally obtained via the following: where the dark current map is first corrected for the dead pixel.","metadata":{"papermill":{"duration":0.012643,"end_time":"2024-08-29T13:25:07.400609","exception":false,"start_time":"2024-08-29T13:25:07.387966","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def clean_dark(signal, dead, dark, dt):\n\n    dark = np.ma.masked_where(dead, dark)\n    dark = np.tile(dark, (signal.shape[0], 1, 1))\n\n    signal -= dark* dt[:, np.newaxis, np.newaxis]\n    return signal\n","metadata":{"papermill":{"duration":0.028125,"end_time":"2024-08-29T13:25:07.441764","exception":false,"start_time":"2024-08-29T13:25:07.413639","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.24167Z","iopub.execute_input":"2025-06-27T08:08:01.242042Z","iopub.status.idle":"2025-06-27T08:08:01.279074Z","shell.execute_reply.started":"2025-06-27T08:08:01.24201Z","shell.execute_reply":"2025-06-27T08:08:01.275459Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 4: Get Correlated Double Sampling (CDS)","metadata":{"papermill":{"duration":0.013727,"end_time":"2024-08-29T13:25:07.47284","exception":false,"start_time":"2024-08-29T13:25:07.459113","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"The science frames are alternating between the start of the exposure and the end of the exposure. The lecture scheme is a ramp with a double sampling, called Correlated Double Sampling (CDS), the detector is read twice, once at the start of the exposure and once at the end of the exposure. The final CDS is the difference (End of exposure) - (Start of exposure).","metadata":{"papermill":{"duration":0.017064,"end_time":"2024-08-29T13:25:07.505488","exception":false,"start_time":"2024-08-29T13:25:07.488424","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def get_cds(signal):\n    cds = signal[:,1::2,:,:] - signal[:,::2,:,:]\n    return cds","metadata":{"papermill":{"duration":0.022926,"end_time":"2024-08-29T13:25:07.541562","exception":false,"start_time":"2024-08-29T13:25:07.518636","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.285072Z","iopub.execute_input":"2025-06-27T08:08:01.285485Z","iopub.status.idle":"2025-06-27T08:08:01.315186Z","shell.execute_reply.started":"2025-06-27T08:08:01.285437Z","shell.execute_reply":"2025-06-27T08:08:01.31378Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 5 (Optional): Time Binning\nThis step is performed mianly to save space. Time series observations are binned together at specified frequency. \n\n","metadata":{"papermill":{"duration":0.012428,"end_time":"2024-08-29T13:25:07.566372","exception":false,"start_time":"2024-08-29T13:25:07.553944","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def bin_obs(cds_signal,binning):\n    cds_transposed = cds_signal.transpose(0,1,3,2)\n    cds_binned = np.zeros((cds_transposed.shape[0], cds_transposed.shape[1]//binning, cds_transposed.shape[2], cds_transposed.shape[3]))\n    for i in range(cds_transposed.shape[1]//binning):\n        cds_binned[:,i,:,:] = np.sum(cds_transposed[:,i*binning:(i+1)*binning,:,:], axis=1)\n    return cds_binned","metadata":{"papermill":{"duration":0.023997,"end_time":"2024-08-29T13:25:07.603229","exception":false,"start_time":"2024-08-29T13:25:07.579232","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.315821Z","iopub.execute_input":"2025-06-27T08:08:01.316027Z","iopub.status.idle":"2025-06-27T08:08:01.364544Z","shell.execute_reply.started":"2025-06-27T08:08:01.316013Z","shell.execute_reply":"2025-06-27T08:08:01.361138Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## Step 6: Flat Field Correction\n","metadata":{"papermill":{"duration":0.015633,"end_time":"2024-08-29T13:25:07.634973","exception":false,"start_time":"2024-08-29T13:25:07.61934","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"The flat field is a map of the detector response to uniform illumination, to correct for the pixel-to-pixel variations of the detector, for example the different quantum efficiencies of each pixel.","metadata":{"papermill":{"duration":0.011976,"end_time":"2024-08-29T13:25:07.660972","exception":false,"start_time":"2024-08-29T13:25:07.648996","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def correct_flat_field(flat,dead, signal):\n    flat = flat.transpose(1, 0)\n    dead = dead.transpose(1, 0)\n    flat = np.ma.masked_where(dead, flat)\n    flat = np.tile(flat, (signal.shape[0], 1, 1))\n    signal = signal / flat\n    return signal","metadata":{"papermill":{"duration":0.025161,"end_time":"2024-08-29T13:25:07.698534","exception":false,"start_time":"2024-08-29T13:25:07.673373","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.366389Z","iopub.execute_input":"2025-06-27T08:08:01.367111Z","iopub.status.idle":"2025-06-27T08:08:01.407711Z","shell.execute_reply.started":"2025-06-27T08:08:01.367075Z","shell.execute_reply":"2025-06-27T08:08:01.404835Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Calibrating all training data","metadata":{"papermill":{"duration":0.012076,"end_time":"2024-08-29T13:25:07.726916","exception":false,"start_time":"2024-08-29T13:25:07.71484","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"you can choose to correct the non-linearity of the pixels' response, to apply flat field, dark and dead map or to leave the data unchanged. The observations are binned in time by group of 30 frames for AIRS and 360 frames for FGS1, to obtain a lighter data-cube, easier to use. The images are cut along the wavelength axis between pixels 39 and 321, so that the 282 pixels left in the wavelength dimension match the last 282 targets' points, from AIRS. The 283rd targets' point is the one for FGS1 that will be added later on. ","metadata":{"papermill":{"duration":0.013,"end_time":"2024-08-29T13:25:07.75262","exception":false,"start_time":"2024-08-29T13:25:07.73962","status":"completed"},"tags":[]}},{"cell_type":"code","source":"## we will start by getting the index of the training data:\ndef get_index(files,CHUNKS_SIZE ):\n    index = []\n    for file in files:\n        file_name = file.split('/')[-1]\n        if file_name.split('_')[0] == 'AIRS-CH0' and file_name.split('_')[1] == 'signal' and file_name.split('_')[2] == '0.parquet':\n            file_index = os.path.basename(os.path.dirname(file))\n            index.append(int(file_index))\n    index = np.array(index)\n    index = np.sort(index) \n    # credit to DennisSakva\n    index=np.array_split(index, len(index)//CHUNKS_SIZE)\n    \n    return index","metadata":{"papermill":{"duration":0.024197,"end_time":"2024-08-29T13:25:07.789483","exception":false,"start_time":"2024-08-29T13:25:07.765286","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.40999Z","iopub.execute_input":"2025-06-27T08:08:01.410391Z","iopub.status.idle":"2025-06-27T08:08:01.442862Z","shell.execute_reply.started":"2025-06-27T08:08:01.410361Z","shell.execute_reply":"2025-06-27T08:08:01.441826Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"files = glob.glob(os.path.join(path_folder + 'train/', '*/*'))\n\nindex = get_index(files[:300],CHUNKS_SIZE)  ## 22 is hardcoded here but please feel free to remove it if you want to do it for the entire dataset\n\naxis_info = pd.read_parquet(os.path.join(path_folder,'axis_info.parquet'))\nDO_MASK = True\nDO_THE_NL_CORR = False\nDO_DARK = True\nDO_FLAT = True\nTIME_BINNING = True\n\ncut_inf, cut_sup = 39, 321\nl = cut_sup - cut_inf\n\nfor n, index_chunk in enumerate(tqdm(index)):\n    AIRS_CH0_clean = np.zeros((CHUNKS_SIZE, 11250, 32, l))\n    FGS1_clean = np.zeros((CHUNKS_SIZE, 135000, 32, 32))\n    \n    for i in range (CHUNKS_SIZE) : \n        df = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/AIRS-CH0_signal_0.parquet'))\n        signal = df.values.astype(np.float64).reshape((df.shape[0], 32, 356))\n\n        signal = ADC_convert(signal,)\n        dt_airs = axis_info['AIRS-CH0-integration_time'].dropna().values\n        dt_airs[1::2] += 0.1\n        chopped_signal = signal[:, :, cut_inf:cut_sup]\n        del signal, df\n        \n        # CLEANING THE DATA: AIRS\n        flat = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/AIRS-CH0_calibration_0/flat.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        dark = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/AIRS-CH0_calibration_0/dark.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        dead_airs = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/AIRS-CH0_calibration_0/dead.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        linear_corr = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/AIRS-CH0_calibration_0/linear_corr.parquet')).values.astype(np.float64).reshape((6, 32, 356))[:, :, cut_inf:cut_sup]\n        \n        if DO_MASK:\n            chopped_signal = mask_hot_dead(chopped_signal, dead_airs, dark)\n            AIRS_CH0_clean[i] = chopped_signal\n        else:\n            AIRS_CH0_clean[i] = chopped_signal\n            \n        if DO_THE_NL_CORR: \n            linear_corr_signal = apply_linear_corr(linear_corr,AIRS_CH0_clean[i])\n            AIRS_CH0_clean[i,:, :, :] = linear_corr_signal\n        del linear_corr\n        \n        if DO_DARK: \n            cleaned_signal = clean_dark(AIRS_CH0_clean[i], dead_airs, dark, dt_airs)\n            AIRS_CH0_clean[i] = cleaned_signal\n        else: \n            pass\n        del dark\n        \n        df = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/FGS1_signal_0.parquet'))\n        fgs_signal = df.values.astype(np.float64).reshape((df.shape[0], 32, 32))\n\n        \n        fgs_signal = ADC_convert(fgs_signal, )\n        dt_fgs1 = np.ones(len(fgs_signal))*0.1\n        dt_fgs1[1::2] += 0.1\n        chopped_FGS1 = fgs_signal\n        del fgs_signal, df\n        \n        # CLEANING THE DATA: FGS1\n        flat = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/FGS1_calibration_0/flat.parquet')).values.astype(np.float64).reshape((32, 32))\n        dark = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/FGS1_calibration_0/dark.parquet')).values.astype(np.float64).reshape((32, 32))\n        dead_fgs1 = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/FGS1_calibration_0/dead.parquet')).values.astype(np.float64).reshape((32, 32))\n        linear_corr = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/FGS1_calibration_0/linear_corr.parquet')).values.astype(np.float64).reshape((6, 32, 32))\n        \n        if DO_MASK:\n            chopped_FGS1 = mask_hot_dead(chopped_FGS1, dead_fgs1, dark)\n            FGS1_clean[i] = chopped_FGS1\n        else:\n            FGS1_clean[i] = chopped_FGS1\n\n        if DO_THE_NL_CORR: \n            linear_corr_signal = apply_linear_corr(linear_corr,FGS1_clean[i])\n            FGS1_clean[i,:, :, :] = linear_corr_signal\n        del linear_corr\n        \n        if DO_DARK: \n            cleaned_signal = clean_dark(FGS1_clean[i], dead_fgs1, dark,dt_fgs1)\n            FGS1_clean[i] = cleaned_signal\n        else: \n            pass\n        del dark\n        \n    # SAVE DATA AND FREE SPACE\n    AIRS_cds = get_cds(AIRS_CH0_clean)\n    FGS1_cds = get_cds(FGS1_clean)\n    \n    del AIRS_CH0_clean, FGS1_clean\n    \n    ## (Optional) Time Binning to reduce space\n    if TIME_BINNING:\n        AIRS_cds_binned = bin_obs(AIRS_cds,binning=30)\n        FGS1_cds_binned = bin_obs(FGS1_cds,binning=30*12)\n    else:\n        AIRS_cds = AIRS_cds.transpose(0,1,3,2) ## this is important to make it consistent for flat fielding, but you can always change it\n        AIRS_cds_binned = AIRS_cds\n        FGS1_cds = FGS1_cds.transpose(0,1,3,2)\n        FGS1_cds_binned = FGS1_cds\n    \n    del AIRS_cds, FGS1_cds\n    \n    for i in range (CHUNKS_SIZE):\n        flat_airs = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/AIRS-CH0_calibration_0/flat.parquet')).values.astype(np.float64).reshape((32, 356))[:, cut_inf:cut_sup]\n        flat_fgs = pd.read_parquet(os.path.join(path_folder,f'train/{index_chunk[i]}/FGS1_calibration_0/flat.parquet')).values.astype(np.float64).reshape((32, 32))\n        if DO_FLAT:\n            corrected_AIRS_cds_binned = correct_flat_field(flat_airs,dead_airs, AIRS_cds_binned[i])\n            AIRS_cds_binned[i] = corrected_AIRS_cds_binned\n            corrected_FGS1_cds_binned = correct_flat_field(flat_fgs,dead_fgs1, FGS1_cds_binned[i])\n            FGS1_cds_binned[i] = corrected_FGS1_cds_binned\n        else:\n            pass\n\n    ## save data\n    np.save(os.path.join(path_out, 'AIRS_clean_train_{}.npy'.format(n)), AIRS_cds_binned)\n    np.save(os.path.join(path_out, 'FGS1_train_{}.npy'.format(n)), FGS1_cds_binned)\n    del AIRS_cds_binned\n    del FGS1_cds_binned","metadata":{"papermill":{"duration":114.333225,"end_time":"2024-08-29T13:27:02.134952","exception":false,"start_time":"2024-08-29T13:25:07.801727","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:08:01.444037Z","iopub.execute_input":"2025-06-27T08:08:01.444408Z","iopub.status.idle":"2025-06-27T08:26:55.736848Z","shell.execute_reply.started":"2025-06-27T08:08:01.444376Z","shell.execute_reply":"2025-06-27T08:26:55.734849Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Once all the chunks are saved, we concatenate them back in a single dataset. This step is simply to save HDD space, modify it as you wish. ","metadata":{"papermill":{"duration":0.014271,"end_time":"2024-08-29T13:27:02.162064","exception":false,"start_time":"2024-08-29T13:27:02.147793","status":"completed"},"tags":[]}},{"cell_type":"code","source":"def load_data (file, chunk_size, nb_files) : \n    data0 = np.load(file + '_0.npy')\n    data_all = np.zeros((nb_files*chunk_size, data0.shape[1], data0.shape[2], data0.shape[3]))\n    data_all[:chunk_size] = data0\n    for i in range (1, nb_files) : \n        data_all[i*chunk_size:(i+1)*chunk_size] = np.load(file + '_{}.npy'.format(i))\n    return data_all \n\ndata_train = load_data(path_out + 'AIRS_clean_train', CHUNKS_SIZE, len(index)) \ndata_train_FGS = load_data(path_out + 'FGS1_train', CHUNKS_SIZE, len(index))\n","metadata":{"papermill":{"duration":0.134081,"end_time":"2024-08-29T13:27:02.309386","exception":false,"start_time":"2024-08-29T13:27:02.175305","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:26:55.739836Z","iopub.execute_input":"2025-06-27T08:26:55.74042Z","iopub.status.idle":"2025-06-27T08:26:57.211291Z","shell.execute_reply.started":"2025-06-27T08:26:55.740369Z","shell.execute_reply":"2025-06-27T08:26:57.210205Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"np.save('./' + 'data_train.npy', data_train)\nnp.save('./' + 'data_train_FGS.npy', data_train_FGS)","metadata":{"papermill":{"duration":0.101539,"end_time":"2024-08-29T13:27:02.425117","exception":false,"start_time":"2024-08-29T13:27:02.323578","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:26:57.212525Z","iopub.execute_input":"2025-06-27T08:26:57.21289Z","iopub.status.idle":"2025-06-27T08:26:57.756373Z","shell.execute_reply.started":"2025-06-27T08:26:57.212866Z","shell.execute_reply":"2025-06-27T08:26:57.755307Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# Plots","metadata":{"papermill":{"duration":0.014312,"end_time":"2024-08-29T13:27:02.452873","exception":false,"start_time":"2024-08-29T13:27:02.438561","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"Let us checks that everything went well during the data import. ","metadata":{"papermill":{"duration":0.012982,"end_time":"2024-08-29T13:27:02.479542","exception":false,"start_time":"2024-08-29T13:27:02.46656","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import matplotlib.pyplot as plt \n\nprint('Shape of the training datasset: \\t')\nprint('\\n For AIRS-CH0:', data_train.shape)\nprint('\\n For FGS1:', data_train_FGS.shape)\n\n","metadata":{"papermill":{"duration":0.02921,"end_time":"2024-08-29T13:27:02.521934","exception":false,"start_time":"2024-08-29T13:27:02.492724","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:26:57.757555Z","iopub.execute_input":"2025-06-27T08:26:57.757973Z","iopub.status.idle":"2025-06-27T08:26:57.764545Z","shell.execute_reply.started":"2025-06-27T08:26:57.757943Z","shell.execute_reply":"2025-06-27T08:26:57.763257Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Plot of some images: ","metadata":{"papermill":{"duration":0.0146,"end_time":"2024-08-29T13:27:02.553564","exception":false,"start_time":"2024-08-29T13:27:02.538964","status":"completed"},"tags":[]}},{"cell_type":"code","source":"plt.imshow(data_train_FGS[-1,50,:,:].T, aspect = 'auto')","metadata":{"papermill":{"duration":0.41396,"end_time":"2024-08-29T13:27:02.981707","exception":false,"start_time":"2024-08-29T13:27:02.567747","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:26:57.765917Z","iopub.execute_input":"2025-06-27T08:26:57.766221Z","iopub.status.idle":"2025-06-27T08:26:58.202481Z","shell.execute_reply.started":"2025-06-27T08:26:57.766199Z","shell.execute_reply":"2025-06-27T08:26:58.201662Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Plot of some light-curves: ","metadata":{"papermill":{"duration":0.083722,"end_time":"2024-08-29T13:27:03.079452","exception":false,"start_time":"2024-08-29T13:27:02.99573","status":"completed"},"tags":[]}},{"cell_type":"code","source":"\nfor i in range(len(data_train)) : \n    light_curve = data_train[i,:,:,:].sum(axis=(1,2))\n    plt.plot(light_curve/light_curve.mean(), '-', alpha=0.3)\n\nplt.xlabel('Time (frame index)')\nplt.ylabel('Normalized flux in the frame')","metadata":{"papermill":{"duration":0.375806,"end_time":"2024-08-29T13:27:03.469794","exception":false,"start_time":"2024-08-29T13:27:03.093988","status":"completed"},"tags":[],"trusted":true,"execution":{"iopub.status.busy":"2025-06-27T08:26:58.203756Z","iopub.execute_input":"2025-06-27T08:26:58.204026Z","iopub.status.idle":"2025-06-27T08:26:58.607094Z","shell.execute_reply.started":"2025-06-27T08:26:58.204002Z","shell.execute_reply":"2025-06-27T08:26:58.606014Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"papermill":{"duration":0.015818,"end_time":"2024-08-29T13:27:03.502235","exception":false,"start_time":"2024-08-29T13:27:03.486417","status":"completed"},"tags":[],"trusted":true},"outputs":[],"execution_count":null}]}