{
  "id": 114426,
  "title": "How are we normalizing the data",
  "url": "/competitions/rsna-intracranial-hemorrhage-detection/discussion/114426",
  "author_name": "Jaideep",
  "post_date": "2019-10-26T05:35:56.144000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all..\nm new entrant to this competition .. what i understand from discussion normalizing the data is bit of tricky task. Would appreciate if some one can point me to the right way of doing it here.</p>",
  "messages": [
    {
      "id": 658698,
      "postDate": "2019-10-26T11:42:51.873Z",
      "content": "<p>That's well explained f.e. here <a href=\"https://www.youtube.com/watch?v=KZld-5W99cI\">https://www.youtube.com/watch?v=KZld-5W99cI</a></p>",
      "rawMarkdown": "That's well explained f.e. here https://www.youtube.com/watch?v=KZld-5W99cI",
      "votes": 1
    },
    {
      "id": 658728,
      "postDate": "2019-10-26T12:48:04.487Z",
      "content": "<p>refer to \n<a href=\"https://www.kaggle.com/akensert/inceptionv3-prev-resnet50-keras-baseline-model\">https://www.kaggle.com/akensert/inceptionv3-prev-resnet50-keras-baseline-model</a></p>\n\n<p>```</p>\n\n<h1>-- dicom -</h1>\n\n<p>def window_image(pixel_array, window_center, window_width, is_normalize=True):\n    image_min = window_center - window_width // 2\n    image_max = window_center + window_width // 2\n    image = np.clip(pixel_array, image_min, image_max)</p>\n\n<pre><code>if is_normalize:\n    image = (image-image_min)/(image_max-image_min)\nreturn image\n</code></pre>\n\n<p>def make_image(dicom):</p>\n\n<pre><code>if (dicom.BitsStored == 12) and (dicom.PixelRepresentation == 0) and (int(dicom.RescaleIntercept) &amp;gt; -100):\n#if 0:\n    # see: https://www.kaggle.com/jhoward/cleaning-the-data-for-rapid-prototyping-fastai\n    p = dicom.pixel_array + 1000\n    p[p&amp;gt;=4096] = p[p&amp;gt;=4096] - 4096\n    dicom.PixelData = p.tobytes()\n    dicom.RescaleIntercept = -1000\n\npixel_array = dicom.pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\nbrain       = window_image(pixel_array, 40,  80)\nsubdural    = window_image(pixel_array, 80, 200)\nsoft_tissue = window_image(pixel_array, 40, 380)\n\nimage = np.dstack([soft_tissue,subdural,brain])\nimage = (image*255).astype(np.uint8)\nreturn image\n</code></pre>\n\n<p>```</p>\n\n<p>in pytorch data loader:</p>\n\n<p>```</p>\n\n<h1>---------------------</h1>\n\n<p>NUM_TEST = 471270//6\nNAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\nLABEL_TO_NAME = {v: k for k, v in NAME_TO_LABEL.items()}\nNUM_CLASS = len(NAME_TO_LABEL)</p>\n\n<h1>corrupted images</h1>\n\n<p>INVALID = [\n    'ID_6431af929'\n]</p>\n\n<p>DATA_DIR = '/root/share/project/kaggle/2019/intracranial_hemorrhage/data'</p>\n\n<p>class RSNADataset(Dataset):\n    def <strong>init</strong>(self, split, csv, folder, mode, augment=None):</p>\n\n<pre><code>    self.split   = split\n    self.csv     = csv\n    self.folder  = folder\n    self.mode    = mode\n    self.augment = augment\n\n    uid=[]\n    dir=[]\n    for f,s in zip(folder,split):\n        s = np.load(DATA_DIR + '/split/%s'%s , allow_pickle=True)\n        uid.append(s)\n        dir.append([f]*len(s))\n\n    self.uid = list(np.concatenate(uid))\n    self.dir = list(np.concatenate(dir))\n\n    df = pd.concat([pd.read_csv(DATA_DIR + '/%s'%f).fillna('') for f in csv])\n    df = df_loc_by_list(df, 'image_id',self.uid)\n    self.df = df\n    self.label = df[list(NAME_TO_LABEL.keys())].values\n    self.num_image = len(df)\n\n    assert(list(df.columns[1:])==list(NAME_TO_LABEL.keys()))\n\n\n\ndef __str__(self):\n    label = self.df[list(NAME_TO_LABEL.keys())].values\n    num_pos = label.sum(0)\n    num_neg = self.num_image - num_pos\n\n    #---\n\n    string  = ''\n    string += '\\tmode    = %s\\n'%self.mode\n    string += '\\tsplit   = %s\\n'%self.split\n    string += '\\tcsv     = %s\\n'%str(self.csv)\n    string += '\\tcsv     = %s\\n'%str(self.folder)\n    string += '\\tnum_image = %8d\\n'%self.num_image\n    string += '\\tlen       = %8d\\n'%len(self)\n    if self.mode == 'train':\n        for c in range(NUM_CLASS):\n            pos = num_pos[c]\n            neg = num_neg[c]\n            num = self.num_image\n            string += '\\t\\t%16s   neg%d, pos%d= %5d  %0.3f,  %5d  %0.3f\\n'%(LABEL_TO_NAME[c], c,c,neg,neg/num,pos,pos/num)\n\n    return string\n\n\ndef __len__(self):\n    return len(self.uid)\n\n\ndef __getitem__(self, index):\n    # print(index)\n\n    dir, image_id = self.dir[index], self.uid[index]\n    #image_id = 'ID_0000aee4b'\n    dicom = pydicom.dcmread(DATA_DIR + '/%s/%s.dcm'%(dir, image_id))\n    image = make_image(dicom)\n    label = self.label[index]\n\n    infor = Struct(\n        index    = index,\n        image_id = image_id,\n    )\n\n    if self.augment is None:\n        return image, label, infor\n    else:\n        return self.augment(image, label, infor)\n</code></pre>\n\n#\n\n<p>def run_check_train_dataset():\n    dataset = RSNADataset(\n        mode    = 'train',\n        csv     = ['stage_1_train.more1.csv',],\n        folder  = ['dicom/stage_1_train',],\n        #split   = ['valid_small_fold0_1396.npy',],\n        split   = ['debug_3000.npy',],\n        augment = None, #\n    )\n    print(dataset)\n    #exit(0)</p>\n\n<pre><code>for n in range(0,len(dataset)):\n    i = n #i = np.random.choice(len(dataset))\n\n    image, label, infor = dataset[i]\n\n    #----\n    print('%05d : %s'%(i, infor.image_id))\n    print('label = %s'%str(label))\n    print('')\n    image_show('image',image)\n    cv2.waitKey(0)\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "refer to \nhttps://www.kaggle.com/akensert/inceptionv3-prev-resnet50-keras-baseline-model\n\n```\n\n\n#-- dicom -################################\n\ndef window_image(pixel_array, window_center, window_width, is_normalize=True):\n    image_min = window_center - window_width // 2\n    image_max = window_center + window_width // 2\n    image = np.clip(pixel_array, image_min, image_max)\n\n    if is_normalize:\n        image = (image-image_min)/(image_max-image_min)\n    return image\n\n\ndef make_image(dicom):\n\n    if (dicom.BitsStored == 12) and (dicom.PixelRepresentation == 0) and (int(dicom.RescaleIntercept) &gt; -100):\n    #if 0:\n        # see: https://www.kaggle.com/jhoward/cleaning-the-data-for-rapid-prototyping-fastai\n        p = dicom.pixel_array + 1000\n        p[p&gt;=4096] = p[p&gt;=4096] - 4096\n        dicom.PixelData = p.tobytes()\n        dicom.RescaleIntercept = -1000\n\n    pixel_array = dicom.pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n    brain       = window_image(pixel_array, 40,  80)\n    subdural    = window_image(pixel_array, 80, 200)\n    soft_tissue = window_image(pixel_array, 40, 380)\n\n    image = np.dstack([soft_tissue,subdural,brain])\n    image = (image*255).astype(np.uint8)\n    return image\n\n\n\n```\n\nin pytorch data loader:\n\n```\n#---------------------\nNUM_TEST = 471270//6\nNAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\nLABEL_TO_NAME = {v: k for k, v in NAME_TO_LABEL.items()}\nNUM_CLASS = len(NAME_TO_LABEL)\n\n# corrupted images\nINVALID = [\n    'ID_6431af929'\n]\n\nDATA_DIR = '/root/share/project/kaggle/2019/intracranial_hemorrhage/data'\n\nclass RSNADataset(Dataset):\n    def __init__(self, split, csv, folder, mode, augment=None):\n\n        self.split   = split\n        self.csv     = csv\n        self.folder  = folder\n        self.mode    = mode\n        self.augment = augment\n\n        uid=[]\n        dir=[]\n        for f,s in zip(folder,split):\n            s = np.load(DATA_DIR + '/split/%s'%s , allow_pickle=True)\n            uid.append(s)\n            dir.append([f]*len(s))\n\n        self.uid = list(np.concatenate(uid))\n        self.dir = list(np.concatenate(dir))\n\n        df = pd.concat([pd.read_csv(DATA_DIR + '/%s'%f).fillna('') for f in csv])\n        df = df_loc_by_list(df, 'image_id',self.uid)\n        self.df = df\n        self.label = df[list(NAME_TO_LABEL.keys())].values\n        self.num_image = len(df)\n\n        assert(list(df.columns[1:])==list(NAME_TO_LABEL.keys()))\n\n\n\n    def __str__(self):\n        label = self.df[list(NAME_TO_LABEL.keys())].values\n        num_pos = label.sum(0)\n        num_neg = self.num_image - num_pos\n\n        #---\n\n        string  = ''\n        string += '\\tmode    = %s\\n'%self.mode\n        string += '\\tsplit   = %s\\n'%self.split\n        string += '\\tcsv     = %s\\n'%str(self.csv)\n        string += '\\tcsv     = %s\\n'%str(self.folder)\n        string += '\\tnum_image = %8d\\n'%self.num_image\n        string += '\\tlen       = %8d\\n'%len(self)\n        if self.mode == 'train':\n            for c in range(NUM_CLASS):\n                pos = num_pos[c]\n                neg = num_neg[c]\n                num = self.num_image\n                string += '\\t\\t%16s   neg%d, pos%d= %5d  %0.3f,  %5d  %0.3f\\n'%(LABEL_TO_NAME[c], c,c,neg,neg/num,pos,pos/num)\n\n        return string\n\n\n    def __len__(self):\n        return len(self.uid)\n\n\n    def __getitem__(self, index):\n        # print(index)\n\n        dir, image_id = self.dir[index], self.uid[index]\n        #image_id = 'ID_0000aee4b'\n        dicom = pydicom.dcmread(DATA_DIR + '/%s/%s.dcm'%(dir, image_id))\n        image = make_image(dicom)\n        label = self.label[index]\n\n        infor = Struct(\n            index    = index,\n            image_id = image_id,\n        )\n\n        if self.augment is None:\n            return image, label, infor\n        else:\n            return self.augment(image, label, infor)\n\n###################################\ndef run_check_train_dataset():\n    dataset = RSNADataset(\n        mode    = 'train',\n        csv     = ['stage_1_train.more1.csv',],\n        folder  = ['dicom/stage_1_train',],\n        #split   = ['valid_small_fold0_1396.npy',],\n        split   = ['debug_3000.npy',],\n        augment = None, #\n    )\n    print(dataset)\n    #exit(0)\n\n    for n in range(0,len(dataset)):\n        i = n #i = np.random.choice(len(dataset))\n\n        image, label, infor = dataset[i]\n\n        #----\n        print('%05d : %s'%(i, infor.image_id))\n        print('label = %s'%str(label))\n        print('')\n        image_show('image',image)\n        cv2.waitKey(0)\n```",
      "votes": 2,
      "replies": [
        {
          "id": 658909,
          "postDate": "2019-10-26T18:10:58.020Z",
          "content": "<p>Thanks, for you code. What is <code>split   = ['debug_3000.npy',]</code> in <code>run_check_train_dataset</code>?</p>",
          "rawMarkdown": "Thanks, for you code. What is `split   = ['debug_3000.npy',]` in `run_check_train_dataset`?"
        },
        {
          "id": 659933,
          "postDate": "2019-10-28T13:18:10.083Z",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Thanks</p>\n\n<p>```\n pixel_array = dicom.pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n    brain       = window_image(pixel_array, 40,  80)\n    subdural    = window_image(pixel_array, 80, 200)\n    soft_tissue = window_image(pixel_array, 40, 380)</p>\n\n<pre><code>image = np.dstack([soft_tissue,subdural,brain])\nimage = (image*255).astype(np.uint8)\nreturn image\n</code></pre>\n\n<p>```\n1) why do we need to apply all three windows,image <br>\n2) I generally see image.div_ (255)  ,here it is multiply  sorry it may be very dumb q :)</p>",
          "rawMarkdown": "@hengck23  Thanks\n\n```\n pixel_array = dicom.pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n    brain       = window_image(pixel_array, 40,  80)\n    subdural    = window_image(pixel_array, 80, 200)\n    soft_tissue = window_image(pixel_array, 40, 380)\n\n    image = np.dstack([soft_tissue,subdural,brain])\n    image = (image*255).astype(np.uint8)\n    return image\n```\n1) why do we need to apply all three windows,image  \n2) I generally see image.div_ (255)  ,here it is multiply  sorry it may be very dumb q :)"
        }
      ]
    },
    {
      "id": 658525,
      "postDate": "2019-10-26T05:35:56.143Z",
      "content": "<p>Hi all..\nm new entrant to this competition .. what i understand from discussion normalizing the data is bit of tricky task. Would appreciate if some one can point me to the right way of doing it here.</p>",
      "rawMarkdown": "Hi all..\nm new entrant to this competition .. what i understand from discussion normalizing the data is bit of tricky task. Would appreciate if some one can point me to the right way of doing it here.\n",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 658698,
      "author_name": "Oleg Yaroshevskiy",
      "author_url": "",
      "post_date": "2019-10-26T11:42:51.873000",
      "content": "<p>That's well explained f.e. here <a href=\"https://www.youtube.com/watch?v=KZld-5W99cI\">https://www.youtube.com/watch?v=KZld-5W99cI</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 658728,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-10-26T12:48:04.487000",
      "content": "<p>refer to \n<a href=\"https://www.kaggle.com/akensert/inceptionv3-prev-resnet50-keras-baseline-model\">https://www.kaggle.com/akensert/inceptionv3-prev-resnet50-keras-baseline-model</a></p>\n\n<p>```</p>\n\n<h1>-- dicom -</h1>\n\n<p>def window_image(pixel_array, window_center, window_width, is_normalize=True):\n    image_min = window_center - window_width // 2\n    image_max = window_center + window_width // 2\n    image = np.clip(pixel_array, image_min, image_max)</p>\n\n<pre><code>if is_normalize:\n    image = (image-image_min)/(image_max-image_min)\nreturn image\n</code></pre>\n\n<p>def make_image(dicom):</p>\n\n<pre><code>if (dicom.BitsStored == 12) and (dicom.PixelRepresentation == 0) and (int(dicom.RescaleIntercept) &amp;gt; -100):\n#if 0:\n    # see: https://www.kaggle.com/jhoward/cleaning-the-data-for-rapid-prototyping-fastai\n    p = dicom.pixel_array + 1000\n    p[p&amp;gt;=4096] = p[p&amp;gt;=4096] - 4096\n    dicom.PixelData = p.tobytes()\n    dicom.RescaleIntercept = -1000\n\npixel_array = dicom.pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\nbrain       = window_image(pixel_array, 40,  80)\nsubdural    = window_image(pixel_array, 80, 200)\nsoft_tissue = window_image(pixel_array, 40, 380)\n\nimage = np.dstack([soft_tissue,subdural,brain])\nimage = (image*255).astype(np.uint8)\nreturn image\n</code></pre>\n\n<p>```</p>\n\n<p>in pytorch data loader:</p>\n\n<p>```</p>\n\n<h1>---------------------</h1>\n\n<p>NUM_TEST = 471270//6\nNAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\nLABEL_TO_NAME = {v: k for k, v in NAME_TO_LABEL.items()}\nNUM_CLASS = len(NAME_TO_LABEL)</p>\n\n<h1>corrupted images</h1>\n\n<p>INVALID = [\n    'ID_6431af929'\n]</p>\n\n<p>DATA_DIR = '/root/share/project/kaggle/2019/intracranial_hemorrhage/data'</p>\n\n<p>class RSNADataset(Dataset):\n    def <strong>init</strong>(self, split, csv, folder, mode, augment=None):</p>\n\n<pre><code>    self.split   = split\n    self.csv     = csv\n    self.folder  = folder\n    self.mode    = mode\n    self.augment = augment\n\n    uid=[]\n    dir=[]\n    for f,s in zip(folder,split):\n        s = np.load(DATA_DIR + '/split/%s'%s , allow_pickle=True)\n        uid.append(s)\n        dir.append([f]*len(s))\n\n    self.uid = list(np.concatenate(uid))\n    self.dir = list(np.concatenate(dir))\n\n    df = pd.concat([pd.read_csv(DATA_DIR + '/%s'%f).fillna('') for f in csv])\n    df = df_loc_by_list(df, 'image_id',self.uid)\n    self.df = df\n    self.label = df[list(NAME_TO_LABEL.keys())].values\n    self.num_image = len(df)\n\n    assert(list(df.columns[1:])==list(NAME_TO_LABEL.keys()))\n\n\n\ndef __str__(self):\n    label = self.df[list(NAME_TO_LABEL.keys())].values\n    num_pos = label.sum(0)\n    num_neg = self.num_image - num_pos\n\n    #---\n\n    string  = ''\n    string += '\\tmode    = %s\\n'%self.mode\n    string += '\\tsplit   = %s\\n'%self.split\n    string += '\\tcsv     = %s\\n'%str(self.csv)\n    string += '\\tcsv     = %s\\n'%str(self.folder)\n    string += '\\tnum_image = %8d\\n'%self.num_image\n    string += '\\tlen       = %8d\\n'%len(self)\n    if self.mode == 'train':\n        for c in range(NUM_CLASS):\n            pos = num_pos[c]\n            neg = num_neg[c]\n            num = self.num_image\n            string += '\\t\\t%16s   neg%d, pos%d= %5d  %0.3f,  %5d  %0.3f\\n'%(LABEL_TO_NAME[c], c,c,neg,neg/num,pos,pos/num)\n\n    return string\n\n\ndef __len__(self):\n    return len(self.uid)\n\n\ndef __getitem__(self, index):\n    # print(index)\n\n    dir, image_id = self.dir[index], self.uid[index]\n    #image_id = 'ID_0000aee4b'\n    dicom = pydicom.dcmread(DATA_DIR + '/%s/%s.dcm'%(dir, image_id))\n    image = make_image(dicom)\n    label = self.label[index]\n\n    infor = Struct(\n        index    = index,\n        image_id = image_id,\n    )\n\n    if self.augment is None:\n        return image, label, infor\n    else:\n        return self.augment(image, label, infor)\n</code></pre>\n\n#\n\n<p>def run_check_train_dataset():\n    dataset = RSNADataset(\n        mode    = 'train',\n        csv     = ['stage_1_train.more1.csv',],\n        folder  = ['dicom/stage_1_train',],\n        #split   = ['valid_small_fold0_1396.npy',],\n        split   = ['debug_3000.npy',],\n        augment = None, #\n    )\n    print(dataset)\n    #exit(0)</p>\n\n<pre><code>for n in range(0,len(dataset)):\n    i = n #i = np.random.choice(len(dataset))\n\n    image, label, infor = dataset[i]\n\n    #----\n    print('%05d : %s'%(i, infor.image_id))\n    print('label = %s'%str(label))\n    print('')\n    image_show('image',image)\n    cv2.waitKey(0)\n</code></pre>\n\n<p>```</p>",
      "votes": 2,
      "replies": [
        {
          "id": 658909,
          "author_name": "Den Petrov",
          "author_url": "",
          "post_date": "2019-10-26T18:10:58.020000",
          "content": "<p>Thanks, for you code. What is <code>split   = ['debug_3000.npy',]</code> in <code>run_check_train_dataset</code>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659933,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-10-28T13:18:10.083000",
          "content": "<p><a href=\"/hengck23\">@hengck23</a>  Thanks</p>\n\n<p>```\n pixel_array = dicom.pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n    brain       = window_image(pixel_array, 40,  80)\n    subdural    = window_image(pixel_array, 80, 200)\n    soft_tissue = window_image(pixel_array, 40, 380)</p>\n\n<pre><code>image = np.dstack([soft_tissue,subdural,brain])\nimage = (image*255).astype(np.uint8)\nreturn image\n</code></pre>\n\n<p>```\n1) why do we need to apply all three windows,image <br>\n2) I generally see image.div_ (255)  ,here it is multiply  sorry it may be very dumb q :)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "658698": "That's well explained f.e. here https://www.youtube.com/watch?v=KZld-5W99cI",
    "658728": "refer to \nhttps://www.kaggle.com/akensert/inceptionv3-prev-resnet50-keras-baseline-model\n\n```\n\n\n#-- dicom -################################\n\ndef window_image(pixel_array, window_center, window_width, is_normalize=True):\n    image_min = window_center - window_width // 2\n    image_max = window_center + window_width // 2\n    image = np.clip(pixel_array, image_min, image_max)\n\n    if is_normalize:\n        image = (image-image_min)/(image_max-image_min)\n    return image\n\n\ndef make_image(dicom):\n\n    if (dicom.BitsStored == 12) and (dicom.PixelRepresentation == 0) and (int(dicom.RescaleIntercept) &gt; -100):\n    #if 0:\n        # see: https://www.kaggle.com/jhoward/cleaning-the-data-for-rapid-prototyping-fastai\n        p = dicom.pixel_array + 1000\n        p[p&gt;=4096] = p[p&gt;=4096] - 4096\n        dicom.PixelData = p.tobytes()\n        dicom.RescaleIntercept = -1000\n\n    pixel_array = dicom.pixel_array * dicom.RescaleSlope + dicom.RescaleIntercept\n    brain       = window_image(pixel_array, 40,  80)\n    subdural    = window_image(pixel_array, 80, 200)\n    soft_tissue = window_image(pixel_array, 40, 380)\n\n    image = np.dstack([soft_tissue,subdural,brain])\n    image = (image*255).astype(np.uint8)\n    return image\n\n\n\n```\n\nin pytorch data loader:\n\n```\n#---------------------\nNUM_TEST = 471270//6\nNAME_TO_LABEL={\n'any' :              0,\n'epidural' :         1,\n'intraparenchymal' : 2,\n'intraventricular' : 3,\n'subarachnoid' :     4,\n'subdural' :         5,\n}\nLABEL_TO_NAME = {v: k for k, v in NAME_TO_LABEL.items()}\nNUM_CLASS = len(NAME_TO_LABEL)\n\n# corrupted images\nINVALID = [\n    'ID_6431af929'\n]\n\nDATA_DIR = '/root/share/project/kaggle/2019/intracranial_hemorrhage/data'\n\nclass RSNADataset(Dataset):\n    def __init__(self, split, csv, folder, mode, augment=None):\n\n        self.split   = split\n        self.csv     = csv\n        self.folder  = folder\n        self.mode    = mode\n        self.augment = augment\n\n        uid=[]\n        dir=[]\n        for f,s in zip(folder,split):\n            s = np.load(DATA_DIR + '/split/%s'%s , allow_pickle=True)\n            uid.append(s)\n            dir.append([f]*len(s))\n\n        self.uid = list(np.concatenate(uid))\n        self.dir = list(np.concatenate(dir))\n\n        df = pd.concat([pd.read_csv(DATA_DIR + '/%s'%f).fillna('') for f in csv])\n        df = df_loc_by_list(df, 'image_id',self.uid)\n        self.df = df\n        self.label = df[list(NAME_TO_LABEL.keys())].values\n        self.num_image = len(df)\n\n        assert(list(df.columns[1:])==list(NAME_TO_LABEL.keys()))\n\n\n\n    def __str__(self):\n        label = self.df[list(NAME_TO_LABEL.keys())].values\n        num_pos = label.sum(0)\n        num_neg = self.num_image - num_pos\n\n        #---\n\n        string  = ''\n        string += '\\tmode    = %s\\n'%self.mode\n        string += '\\tsplit   = %s\\n'%self.split\n        string += '\\tcsv     = %s\\n'%str(self.csv)\n        string += '\\tcsv     = %s\\n'%str(self.folder)\n        string += '\\tnum_image = %8d\\n'%self.num_image\n        string += '\\tlen       = %8d\\n'%len(self)\n        if self.mode == 'train':\n            for c in range(NUM_CLASS):\n                pos = num_pos[c]\n                neg = num_neg[c]\n                num = self.num_image\n                string += '\\t\\t%16s   neg%d, pos%d= %5d  %0.3f,  %5d  %0.3f\\n'%(LABEL_TO_NAME[c], c,c,neg,neg/num,pos,pos/num)\n\n        return string\n\n\n    def __len__(self):\n        return len(self.uid)\n\n\n    def __getitem__(self, index):\n        # print(index)\n\n        dir, image_id = self.dir[index], self.uid[index]\n        #image_id = 'ID_0000aee4b'\n        dicom = pydicom.dcmread(DATA_DIR + '/%s/%s.dcm'%(dir, image_id))\n        image = make_image(dicom)\n        label = self.label[index]\n\n        infor = Struct(\n            index    = index,\n            image_id = image_id,\n        )\n\n        if self.augment is None:\n            return image, label, infor\n        else:\n            return self.augment(image, label, infor)\n\n###################################\ndef run_check_train_dataset():\n    dataset = RSNADataset(\n        mode    = 'train',\n        csv     = ['stage_1_train.more1.csv',],\n        folder  = ['dicom/stage_1_train',],\n        #split   = ['valid_small_fold0_1396.npy',],\n        split   = ['debug_3000.npy',],\n        augment = None, #\n    )\n    print(dataset)\n    #exit(0)\n\n    for n in range(0,len(dataset)):\n        i = n #i = np.random.choice(len(dataset))\n\n        image, label, infor = dataset[i]\n\n        #----\n        print('%05d : %s'%(i, infor.image_id))\n        print('label = %s'%str(label))\n        print('')\n        image_show('image',image)\n        cv2.waitKey(0)\n```",
    "658525": "Hi all..\nm new entrant to this competition .. what i understand from discussion normalizing the data is bit of tricky task. Would appreciate if some one can point me to the right way of doing it here.\n"
  }
}