{"metadata":{"kernelspec":{"display_name":"Python 3 (ipykernel)","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.9.13"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Effortless Data Transfer: Fetching Selective DICOM Images from Kaggle to GCP Storage Bucket Directly\n\n### We will work on Kaggle competetion RSNA Intracranial Hemorrhage Detection dataset\n\nIn this tutorial, we will explore a creative method for transferring the whole dataset or specific DICOM images from the **RSNA Intracranial Hemorrhage Detection** dataset directly to your Google Cloud Platform (GCP) bucket. You can handle the challenges of downloading the entire dataset, which consumes substantial storage space and precious time. Instead, we will see an efficient technique that allows seamless data transfer to your GCP bucket, saving you valuable resources and streamlining your workflow. Here I'm working on Kaggle notebook environment. So Let's dive right in and revolutionize the way your work with data.\n\n- Begin by creating a new Kaggle Notebook. Navigate to the **<> Notebooks** section in the left sidebar and click on the **+ New Notebook** option.\n\n- Now, let's add a reference to the Kaggle data in your Notebook. To do this, locate and click on the **Add data** button, which can be found in the upper-right corner.\n\n- A window will appear. Inside it, switch to the Competition Data tab. Here, find the **RSNA Intracranial Hemorrhage Detection dataset** and click the **Add button** next to it.\n\n- To facilitate access to your Google Cloud Storage service, you'll want to grant Kaggle the necessary permissions. To start, open the **Add-ons** menu and select **Google Cloud Services.**\n\n- Once the window pops up, make sure to select **Cloud Storage.** Finally, click the **Link Account** button to establish the connection.\n\n\nNow install some Python packages if required.","metadata":{}},{"cell_type":"code","source":"!pip install --upgrade pip # (optional)\n!pip install google-cloud-storage\n!pip install pydicom","metadata":{"execution":{"iopub.execute_input":"2023-09-02T06:22:06.925759Z","iopub.status.busy":"2023-09-02T06:22:06.925145Z","iopub.status.idle":"2023-09-02T06:23:18.685328Z","shell.execute_reply":"2023-09-02T06:23:18.68405Z","shell.execute_reply.started":"2023-09-02T06:22:06.925714Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Importing required Python libraries ","metadata":{}},{"cell_type":"code","source":"from google.cloud import storage\nimport pandas as pd\nimport numpy as np\nimport os\nimport pydicom","metadata":{"execution":{"iopub.execute_input":"2023-09-02T06:24:07.248655Z","iopub.status.busy":"2023-09-02T06:24:07.248159Z","iopub.status.idle":"2023-09-02T06:24:07.4091Z","shell.execute_reply":"2023-09-02T06:24:07.408028Z","shell.execute_reply.started":"2023-09-02T06:24:07.248617Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So here I have uploaded **cleaned_train.csv** file that contains **IDs** of all the DICOM images that I want to store into my Storage bucket. I did some Cleaning on the actual CSV file **stage_2_train.csv.** You can also do the same or you can copy the IDs from the actual file into python list. \nRead the CSV file into dataframe.","metadata":{}},{"cell_type":"code","source":"df_cleaned = pd.read_csv('/kaggle/input/fyp-rsna-dataset/cleaned_train.csv')\ndf_cleaned.head()","metadata":{"execution":{"iopub.execute_input":"2023-09-02T06:24:11.649634Z","iopub.status.busy":"2023-09-02T06:24:11.649255Z","iopub.status.idle":"2023-09-02T06:24:12.518366Z","shell.execute_reply":"2023-09-02T06:24:12.517369Z","shell.execute_reply.started":"2023-09-02T06:24:11.649603Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now copy the column ID into python list.","metadata":{}},{"cell_type":"code","source":"IDs_to_download = df_cleaned['ID'].tolist()","metadata":{"execution":{"iopub.execute_input":"2023-09-02T06:24:25.89151Z","iopub.status.busy":"2023-09-02T06:24:25.891077Z","iopub.status.idle":"2023-09-02T06:24:25.920794Z","shell.execute_reply":"2023-09-02T06:24:25.919759Z","shell.execute_reply.started":"2023-09-02T06:24:25.891449Z"}},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Replace **{your-bucket-name}** with your actual bucket name, **{folder-name}** with your actual folder name and **{your-GCP-project-ID}** with your actual project ID.","metadata":{}},{"cell_type":"code","source":"dicom_folder_path = '/kaggle/input/rsna-intracranial-hemorrhage-detection/rsna-intracranial-hemorrhage-detection/stage_2_train/'","metadata":{"execution":{"iopub.execute_input":"2023-09-02T06:24:27.700048Z","iopub.status.busy":"2023-09-02T06:24:27.699678Z","iopub.status.idle":"2023-09-02T06:24:27.705021Z","shell.execute_reply":"2023-09-02T06:24:27.704033Z","shell.execute_reply.started":"2023-09-02T06:24:27.700018Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"bucket_name = 'your-bucket-name'\ndestination_folder = 'folder-name'\nclient = storage.Client(project='your-GCP-project-ID')\nbucket = client.get_bucket(bucket_name)\nfor image_id in IDs_to_download:\n    image_filename = f\"{image_id}.dcm\"\n    image_filepath = os.path.join(dicom_folder_path, image_filename)\n    \n    if os.path.exists(image_filepath):\n        dicom_file = pydicom.dcmread(image_filepath)\n        \n        destination_blob_name = os.path.join(destination_folder, image_filename)\n        blob = bucket.blob(destination_blob_name)\n        blob.upload_from_filename(image_filepath, content_type='application/dicom')\n\nprint(\"Images uploaded successfully\")","metadata":{"execution":{"iopub.execute_input":"2023-09-02T06:29:36.980283Z","iopub.status.busy":"2023-09-02T06:29:36.979911Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}],"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}}