{
  "id": 372567,
  "title": "Image Classification Tips & Tricks",
  "url": "/competitions/rsna-breast-cancer-detection/discussion/372567",
  "author_name": "The Devastator",
  "post_date": "2022-12-16T16:55:04.281000",
  "votes": 188,
  "comment_count": 22,
  "views": 0,
  "content": "<p>The following are image classification tips and tricks you might find useful for this competition. </p>\n<p>Enjoy! </p>\n<hr>\n<h2>Image Classification Tips &amp; Tricks</h2>\n<h4>Augmentations</h4>\n<h4>Color Augmentations</h4>\n<ul>\n<li><strong>Color Skew:</strong><br>\nThis augmentation randomly adjusts the hue, saturation, and brightness of the image by multiplying each channel by a randomly chosen coefficient. The coefficients are chosen from a range of [0:6;1:4] to ensure that the resulting image is not too distorted.</li>\n</ul>\n<pre><code> ():\n    h, s, v = cv2.split(image)\n    h = h * np.random.uniform(low=, high=)\n    s = s * np.random.uniform(low=, high=)\n    v = v * np.random.uniform(low=, high=)\n     cv2.merge((h, s, v))\n</code></pre>\n<ul>\n<li><strong>RGB Norm:</strong><br>\nThis augmentation normalizes the RGB channels of the image by subtracting the mean value of each channel from the values in that channel and dividing by the standard deviation of the channel. This helps to standardize the values in the image and can improve the performance of the model.</li>\n</ul>\n<pre><code> ():\n    r, g, b = cv2.split(image)\n    r = (r - np.mean(r)) / np.std(r)\n    g = (g - np.mean(g)) / np.std(g)\n    b = (b - np.mean(b)) / np.std(b)\n     cv2.merge((r, g, b))\n</code></pre>\n<ul>\n<li><strong>Black and White:</strong><br>\nThis augmentation converts the image to black and white by converting it to the grayscale color space.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n</code></pre>\n<ul>\n<li><strong>Ben Graham: Greyscale + Gaussian Blur:</strong><br>\nThis augmentation converts the image to grayscale and applies a Gaussian blur to smooth out any noise or details in the image.</li>\n</ul>\n<pre><code> ():\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n    image = cv2.GaussianBlur(image, (, ), )\n     image\n</code></pre>\n<ul>\n<li><strong>Hue, Saturation, Brightness:</strong><br>\nThis augmentation converts the image to the HLS color space, which separates the image into its hue, saturation, and brightness channels.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2HLS)\n</code></pre>\n<ul>\n<li><strong>LUV Color Space:</strong><br>\nThis augmentation converts the image to the LUV color space, which is designed to be perceptually uniform and enables more accurate color comparison.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2LUV)\n</code></pre>\n<ul>\n<li><strong>Alpha Channel:</strong><br>\nThis augmentation adds an alpha channel to the image, which can be used for transparency effects.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2RGBA)\n</code></pre>\n<ul>\n<li><strong>YZ Color Space:</strong><br>\nThis augmentation converts the image to the XYZ color space, which is a device-independent color space that allows for more accurate color representation.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2XYZ)\n</code></pre>\n<ul>\n<li><strong>Luma Chroma:</strong><br>\nThis augmentation converts the image to the YCrCb color space, which separates the image into its luma (brightness) and chroma (color) channels.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2YCrCb)\n</code></pre>\n<ul>\n<li><strong>CIE Lab:</strong></li>\n<li>This augmentation converts the image to the CIE Lab color space, which is designed to be perceptually uniform and enables more accurate color comparison.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2Lab)\n</code></pre>\n<ul>\n<li><strong>YUV Color Space:</strong><br>\nThis augmentation converts the image to the YUV color space, which separates the image into its luminance (brightness) and chrominance (color) channels.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2YUV)\n</code></pre>\n<ul>\n<li><strong>Center Crop:</strong><br>\nThis augmentation randomly crops a rectangular area with an aspect ratio of [3/4,4/3], then randomly scales the crop by a factor between [8%,100%], and finally resizes the crop to an img_sizeXimg_size square. This is done randomly on each batch.</li>\n</ul>\n<pre><code>transforms.CenterCrop((, ))\n</code></pre>\n<p><strong>Flippings:</strong><br>\nThis augmentation adds a probability of random horizontal flip to the image. For example, with a probability of 0.5, the image has a 50% chance of being horizontally flipped.</p>\n<pre><code> ():\n     np.random.uniform() &lt; :\n        image = cv2.flip(image, )\n     image\n</code></pre>\n<ul>\n<li><strong>Random Crop:</strong><br>\nThis augmentation randomly crops a rectangular area from the image.</li>\n</ul>\n<pre><code>transforms.RandomCrop((, ))\n</code></pre>\n<ul>\n<li><strong>Random Resized Crop:</strong><br>\nThis augmentation randomly resizes and crops a rectangular area from the image.</li>\n</ul>\n<pre><code>transforms.RandomResizedCrop((, ))\n</code></pre>\n<ul>\n<li><strong>Color Jitter:</strong><br>\nThis augmentation randomly adjusts the brightness, contrast, saturation, and hue of the image.</li>\n</ul>\n<pre><code>transforms.ColorJitter(brightness=, contrast=, saturation=, hue=)\n</code></pre>\n<ul>\n<li><strong>Random Affine:</strong><br>\nThis augmentation randomly applies an affine transformation to the image, which includes rotation, scaling, and shearing.</li>\n</ul>\n<pre><code>transforms.RandomAffine(degrees=, translate=(, ), scale=(, ), shear=)\n</code></pre>\n<ul>\n<li><strong>Random Horizontal Flip:</strong><br>\nrandomly flips the image horizontally with a probability of 0.5.</li>\n</ul>\n<pre><code>transforms.RandomHorizontalFlip()\n</code></pre>\n<ul>\n<li><strong>Random Vertical Flip:</strong><br>\nThis augmentation randomly flips the image vertically with a probability of 0.5.</li>\n</ul>\n<pre><code>transforms.RandomVerticalFlip()\n</code></pre>\n<ul>\n<li><strong>Random Perspective:</strong><br>\nThis augmentation randomly applies a perspective transformation to the image.</li>\n</ul>\n<pre><code>transforms.RandomPerspective()\n</code></pre>\n<ul>\n<li><strong>Random Rotation:</strong><br>\nThis augmentation randomly rotates the image by a given degree range.</li>\n</ul>\n<pre><code>transforms.RandomRotation(degrees=)\n</code></pre>\n<ul>\n<li><strong>Random Invert:</strong><br>\nThis augmentation randomly inverts the colors of the image.</li>\n</ul>\n<pre><code>transforms.RandomInvert()\n</code></pre>\n<ul>\n<li><strong>Random Posterize:</strong><br>\nThis augmentation randomly reduces the number of bits used to represent each pixel value, resulting in a posterized effect.</li>\n</ul>\n<pre><code>transforms.RandomPosterize(bits=)\n</code></pre>\n<ul>\n<li><strong>Random Solarize:</strong><br>\nThis augmentation randomly applies a solarize effect to the image, where pixels above a certain intensity threshold are inverted.</li>\n</ul>\n<pre><code>transforms.RandomSolarize(threshold=)\n</code></pre>\n<ul>\n<li><strong>Random Autocontrast:</strong><br>\nThis augmentation randomly adjusts the contrast of the image by stretching the intensity values to the full available range.</li>\n</ul>\n<pre><code>transforms.RandomAutocontrast()\n</code></pre>\n<ul>\n<li><strong>Random Equalize:</strong><br>\nThis augmentation randomly equalizes the histogram of the image, resulting in increased contrast.</li>\n</ul>\n<pre><code>transforms.RandomEqualize()\n</code></pre>\n<h4>Advanced Augmentations</h4>\n<ul>\n<li><strong>Auto Augment:</strong><br>\nAuto Augment is an augmentation method that uses reinforcement learning to search for the optimal augmentation policies for a given dataset. It has been shown to improve the performance of image classification models.</li>\n</ul>\n<pre><code> autoaugment  AutoAugment\n\nauto_augment = AutoAugment()\nimage = auto_augment(image)\n</code></pre>\n<ul>\n<li><strong>Fast Autoaugment:</strong><br>\nFast Autoaugment is a faster implementation of the Auto Augment method. It uses a neural network to predict the optimal augmentation policies for a given dataset.</li>\n</ul>\n<pre><code> fast_autoaugment  FastAutoAugment\n\nfast_auto_augment = FastAutoAugment()\nimage = fast_auto_augment(image)\n</code></pre>\n<ul>\n<li><strong>Augmix:</strong><br>\nAugmix is an augmentation method that combines multiple augmented images to create a single, more diverse and realistic image. It has been shown to improve the robustness and generalization of image classification models.</li>\n</ul>\n<pre><code> augmix  AugMix\n\naug_mix = AugMix()\nimage = aug_mix(image)\n</code></pre>\n<ul>\n<li><strong>Mixup/Cutout:</strong><br>\nMixup is an augmentation method that combines two images by linearly interpolating their pixel values. Cutout is an augmentation method that randomly removes a rectangular region from the image. These methods have been shown to improve the robustness and generalization of image classification models.</li>\n</ul>\n<blockquote>\n  <ul>\n  <li>\"You take a picture of a cat and add some \"transparent dog\" on top of it. The amount of transparency is a hyperparam.\"</li>\n  </ul>\n</blockquote>\n<pre><code>x=*x1+(-)x2\ny=*x1+(-)y2\n</code></pre>\n<h3>Validation Time Augmentations</h3>\n<ul>\n<li>Adding validation set data augmentation will improve your model performance on test time augmentation and thus increase the absolution performance but if you seek to gain better performance use some of your validation set for data augmentation experiments or data augmentation experiments during training and use another part to increase your performance parameters, this will reduce your model overfitting.</li>\n<li>You can use the test time augmentation by running the same model multiple times.</li>\n<li>Run many validation data augmentation experiments at the same time to find the ones with the greatest impact your model performance.</li>\n<li><strong>USE EXPERIMENTS:</strong> Experiment. Experiment. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. </li>\n</ul>\n<h3>Test Time Augmentations</h3>\n<p>Augmentations can be useful not juse during training but also during test time. <br>\nSimply apply them on prediction and avg the results. </p>\n<ul>\n<li><strong>Predict on different img scales:</strong> Try to predict the same image but on different scales. </li>\n<li><strong>Color Standardization:</strong> Standardizatize as above.</li>\n</ul>\n<h3>Current High Scoring Models</h3>\n<p>Here is a summary of the current highest scoring models on this competition:</p>\n<p><strong>tf_efficientnetv2_s</strong><br>\nUsed on <a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">Theo Viel</a></p>\n<p><strong>seresnext50_32x4d</strong></p>\n<ul>\n<li>Find it on <a href=\"https://www.kaggle.com/code/vslaykovsky/infer-pytorch-aux-targets-weighted-loss-thres\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy\n</a></li>\n<li>Also used on <a href=\"https://www.kaggle.com/code/hengck23/notebooke04a738685\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">Heng</a></li>\n</ul>\n<p><strong>tf_effv2_s_208_402</strong></p>\n<ul>\n<li>Find it on <a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a></li>\n</ul>\n<h3>Models To Consider</h3>\n<p>And some ideas and models to try out.</p>\n<p><strong>Swin Transformer</strong></p>\n<ul>\n<li>Was the SOTA before EfficientNetV2, You can find example of it <a href=\"https://www.kaggle.com/phalanx/train-swin-t-pytorch-lightning/notebook\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>BeIT Transformer</strong></p>\n<ul>\n<li>Also known to be a very strong model. </li>\n</ul>\n<p><strong>ViT Transformers</strong></p>\n<ul>\n<li>This is the first vision transformer released. Still very powerful to this day.</li>\n</ul>\n<p><strong>EfficientNet L2</strong></p>\n<ul>\n<li>The largest efficientnet V1, If you can somehow get it to work in terms of compute power it might yield high performance.</li>\n</ul>\n<h3>Hidden Layers After Your Backbone</h3>\n<p>Adding more layers can be beneficial as you can use them to learn more advanced features, but it can also temper the fine-tuning of your large pretrained model.</p>\n<h3>Unfreeze Layer By Layer</h3>\n<ul>\n<li>A simple trick that can get you a tiny improvement would be to unfreeze layers of your pretrained backbone as the training progress. </li>\n<li><strong>Adding more layers and freezing everything else:</strong> As it turns out: Many solutions could be even further improved by incorporating another training phase after the pretrained model had been trained! This can be done by freezing the pretrained model and adding dense layers after it. You will get a tiny uplift in the upmost leaderboard with that change though. </li>\n</ul>\n<p><strong>Weight freezing and unfreezing in PyTorch</strong></p>\n<pre><code>\n param  model.parameters():\n  param.requires_grad =  \n</code></pre>\n<pre><code>\n param  model.parameters():\n  param.requires_grad =  \n</code></pre>\n<p><strong>Weight freezing and unfreezing in TensorFlow</strong></p>\n<pre><code>\nlayer.trainable = \n</code></pre>\n<pre><code>\nlayer.trainable = \n</code></pre>\n<h3>Learning-rates &amp; LR Shedulers</h3>\n<p>Learning rates and learning rate schedulers affect your models' training performance. Changing the learning rate can have a great impact on performance and also on training convergence. </p>\n<h4>Learning rate schedulers</h4>\n<p>In recent times, One Cycle Cosine schedule has shown to provide better results on multiple NLP tasks, you can use it this way:</p>\n<p><strong>One Cycle Cosine scheduling in PyTorch</strong></p>\n<pre><code> torch.optim.lr_scheduler  CosineAnnealingLR\noptimizer = torch.optim.AdamW(optimizer_grouped_parameters, lr=args.learning_rate, eps=args.adam_epsilon)\nscheduler = CosineAnnealingLR(optimizer, T_max=num_train_optimization_steps)\nnum_training_steps = num_train_optimization_steps / args.gradient_accumulation_steps\n</code></pre>\n<pre><code>\nscheduler.step()\n\n</code></pre>\n<p><strong>One Cycle Cosine scheduling in TensorFlow</strong></p>\n<pre><code>optimizer = tf.keras.optimizers.Adam(learning_rate)\nscheduler = tf.keras.optimizers.schedules.CosineDecay(learning_rate, decay_steps=num_training_steps)\n</code></pre>\n<h4>Some Learning Rate Schedulers Tips</h4>\n<ul>\n<li>Using “Triangular” or “One Cyclic” methods for learning rate scheduling can provide subtle but significant improvements - those intelligent methods of learning rate scheduling are can overcome some of the batch size issues and this will be obvious when using them, especially when plugging them into the multiple learning rate system. </li>\n<li>Take the time to research for the best learning rate scheduling method for your task and the models you are using, it is a very important part of how your model will converge. </li>\n<li>Learning rate schedules can be used to train models with lower batch sizes or multiple learning rates or any combination of those. </li>\n<li>Try low learning rates first to see if a dramatic learning rate increase will help or hurt your performance. </li>\n<li>Increasing the learning rate at later stages of training or multiple learning rates or high batch size or gradient accumulation or learning rate schedulers sometimes will help your model converge better, it is an advanced technique as sometimes it can hurt the performance but only if you give it too large a value - remember to test it. </li>\n<li>Loss scaling can help reduce loss variance and improve gradient flow when using gradient accumulation or multiple learning rates or high batch size, but if you are trying to solve that problem by increasing batch size, try to increase the learning rate instead as it can sometimes yield better performance. </li>\n<li>In recent times, one cycle learning rate scheduling approaches have taken the leading positions on public leaderboards and past winning solutions. </li>\n</ul>\n<h4>Hyperparameters Of Optimizers</h4>\n<blockquote>\n  <p>Crossposting from \"Transformers Hyperparams Tips\"</p>\n</blockquote>\n<p>There are several things you need to know if you seek to get the best performance out of Adam optimizer:</p>\n<ul>\n<li>Finding the best weight decay value can be tricky and experiments (and luck) will be your best friends here. </li>\n<li>Another important hyperparameter is the <code>beta1</code> and <code>beta2</code> used in the Adam optimizer, choosing the best values depends on your task and the data. many new tasks can benefit from lower beta1 and higher beta2 while on established tasks they would perform the opposite. So again: Experiments will be your best friend here.</li>\n<li>In the world of Adam optimizer, the number one rule is not to underestimate the importance of the optimizer epsilon value. The same principle of finding the best weight decay hyperparameter applies here.</li>\n<li>Do not overuse the gradient clipping norm - it might sometimes be helpful when your gradients explode while the opposite is also true - it can prevent convergence on some tasks.</li>\n<li>Gradient accumulation still sows to provide some subtle benefits, I usually accumulate gradients for about 2 steps but if your GPU does not run out of memory, you can push up to 8 gradient accumulation steps. Gradient accumulation is also useful when using mixed-precision. </li>\n</ul>\n<h3>Overfitting and regularization</h3>\n<ul>\n<li>Use dropout! adding layers dropouts between layers usually yields more training stability and more robust results, use dropouts with your hidden layers.</li>\n<li>Dropout can also be used to increase performance by small margins, experiment with setting layer dropouts before training. the task and the model.</li>\n<li>**If you are into regularization: **Regularization can provide marge uplift on performance when your NN is overfitting or underfitting, for normal Machine learning models, L1 or L2 regularization are OK but also try to use additive and hidden layers dropout.</li>\n<li><strong>ALWAYS TEST OUT IDEAS USING EXPERIMENTS:</strong> Use experiments. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. </li>\n<li><strong>Multi Validations:</strong> You can increase your model's robustness to overfitting by using multiple validations. This however comes at the cost of compute time.</li>\n</ul>\n<h3>Optimizer</h3>\n<p>Everyone today are using Adam or AdamW. BUT something to consider is often times if you hyperparam-search SGD with momentum enough, you probably could get better results with it. But again this requires heavy tuning.</p>\n<p><strong>There are several notable optimizers that are important to know about:</strong></p>\n<ul>\n<li><strong>AdamW:</strong>  This is an extension to the Adam algorithm that prevents exponential weights decay of the model's weights in the outer layers as well as encourages penalized hyper volume below the default weights. </li>\n<li><strong>Adafactor:</strong> It was designed to have low memory usage and is scalable. This optimizer can offer significant optimizer performances using multiple GPUs (see below). </li>\n<li><strong>Novograd:</strong> Basically another Adam-like optimizer but with very better properties. It is one of the optimizers used by  to train the bert-large model. </li>\n<li><strong>Ranger:</strong> Ranger optimizer is a very interesting optimizer that has pretty good results in winning solutions it terms of performance optimization but it is obviously not very well known or supported so I leave that up to you to research. </li>\n<li><strong>Lamb:</strong> GPU optimized reusable Adam optimizer developed by the GLUE and QQP competition winner. </li>\n<li><strong>Lookahead:</strong> A popular optimizer you can use on top of other optimizers and it will provide you with some performance gains. </li>\n</ul>\n<h3>Label Smoothing</h3>\n<blockquote>\n  <p>Original Paper: <a href=\"https://arxiv.org/pdf/1906.02629.pdf\" target=\"_blank\">https://arxiv.org/pdf/1906.02629.pdf</a></p>\n</blockquote>\n<p>Basiclly a simple trick: y_true = y_true * (1.0 - ε) + 0.5 * ε [example: ε = 0.001]<br>\nUsually works well. You can find it on some of the high scoring public notebook of this competition.</p>\n<p><strong>Tensorflow:</strong> </p>\n<pre><code>loss = BinaryCrossentropy(label_smoothing = label_smoothing)\n</code></pre>\n<p><strong>Pytorch:</strong></p>\n<pre><code> torch.nn.modules.loss  _WeightedLoss\n\n ():\n     ():\n        ().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n        self.pos_weight = pos_weight\n\n\n     ():\n          &lt;= smoothing &lt; \n         torch.no_grad(): targets = targets * ( - smoothing) +  * smoothing\n         targets\n\n     ():\n        targets = SmoothBCEwLogits._smooth(targets, inputs.size(-), self.smoothing)\n        loss = F.binary_cross_entropy_with_logits(inputs, targets,self.weight, pos_weight = self.pos_weight)\n          self.reduction == : loss = loss.()\n          self.reduction == : loss = loss.mean()\n         loss\n</code></pre>\n<pre><code>    \n</code></pre>\n<h2>Advanced Tricks</h2>\n<h3>Knowledge Distillation</h3>\n<blockquote>\n  <p>Use a large teacher network to guide the learning of a small network.</p>\n</blockquote>\n<ul>\n<li><strong>Train the large model:</strong> Train a large model on the data.</li>\n<li><strong>Calculate soft target:</strong> Use the trained large model to calculate soft target. That is, the output of the softmax after the large model is \"softened\"</li>\n<li><strong>Studnet Model Training:</strong> Train a student model based on the teachers output as an extra soft target loss function on the basis of the large model, and adjust the proportion of the two loss functions through interpulation.</li>\n</ul>\n<h3>Pseudo Labeling</h3>\n<blockquote>\n  <p>Use a model to label unlabeled data (For example test data) then use the new labeled data you have for training your models.</p>\n</blockquote>\n<ul>\n<li><strong>Train the teacher model:</strong> Train a model on the data you have.</li>\n<li><strong>Calculate soft target:</strong> Use the trained large model to calculate soft target for unlabeled data.</li>\n<li><strong>Important:</strong> Use only the targets your model is \"sure\" about** Use only the highest confidence of the OOF predictions as pseudo labels to avoid mistakes as much as possible. (It might not work if you don't do this.).</li>\n<li><strong>Studnet Model Training:</strong> Train a student model on the new labeled data that you have.</li>\n</ul>\n<blockquote>\n  <p><strong>How do you CV pseudo labeling?</strong> Simple: You concat the unlabeled data to your OOF and simply use them just as you would have used stacking. </p>\n</blockquote>\n<h3>Error Analysis</h3>\n<p>An important practice that can save you a lot of time is to use your model for finding harder or broken data samples.<br>\nThere can be many reasons for an image to be \"harder\" for your model, for example, small target objects, different coloring, cut off target, invalid annotations and more.</p>\n<p><strong>Mistakes are Good News!</strong></p>\n<p>These are exactly the samples that separate the top of the leaderboard from the rest of the participants.<br>\nIf you have a hard time explaining what is going on with your model, it might be a good idea to take a look at the validation samples your model struggles with.</p>\n<h3>Finding Your Model's Errors</h3>\n<p>The easiest way to find errors is to sort validation samples by the model's confidence score and see which ones are being predicted with the lowest confidence.</p>\n<pre><code>mistakes_idx = [img_idx  img_idx  ((train))  (pred[img_idx] &gt; ) != target[img_idx]]\nmistakes_preds = pred[mistakes_idx]\nsorted_idx = np.argsort(mistakes_preds)[:]\n\n</code></pre>",
  "messages": [
    {
      "id": 2067389,
      "postDate": "2022-12-16T16:55:04.280Z",
      "content": "<p>The following are image classification tips and tricks you might find useful for this competition. </p>\n<p>Enjoy! </p>\n<hr>\n<h2>Image Classification Tips &amp; Tricks</h2>\n<h4>Augmentations</h4>\n<h4>Color Augmentations</h4>\n<ul>\n<li><strong>Color Skew:</strong><br>\nThis augmentation randomly adjusts the hue, saturation, and brightness of the image by multiplying each channel by a randomly chosen coefficient. The coefficients are chosen from a range of [0:6;1:4] to ensure that the resulting image is not too distorted.</li>\n</ul>\n<pre><code> ():\n    h, s, v = cv2.split(image)\n    h = h * np.random.uniform(low=, high=)\n    s = s * np.random.uniform(low=, high=)\n    v = v * np.random.uniform(low=, high=)\n     cv2.merge((h, s, v))\n</code></pre>\n<ul>\n<li><strong>RGB Norm:</strong><br>\nThis augmentation normalizes the RGB channels of the image by subtracting the mean value of each channel from the values in that channel and dividing by the standard deviation of the channel. This helps to standardize the values in the image and can improve the performance of the model.</li>\n</ul>\n<pre><code> ():\n    r, g, b = cv2.split(image)\n    r = (r - np.mean(r)) / np.std(r)\n    g = (g - np.mean(g)) / np.std(g)\n    b = (b - np.mean(b)) / np.std(b)\n     cv2.merge((r, g, b))\n</code></pre>\n<ul>\n<li><strong>Black and White:</strong><br>\nThis augmentation converts the image to black and white by converting it to the grayscale color space.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n</code></pre>\n<ul>\n<li><strong>Ben Graham: Greyscale + Gaussian Blur:</strong><br>\nThis augmentation converts the image to grayscale and applies a Gaussian blur to smooth out any noise or details in the image.</li>\n</ul>\n<pre><code> ():\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n    image = cv2.GaussianBlur(image, (, ), )\n     image\n</code></pre>\n<ul>\n<li><strong>Hue, Saturation, Brightness:</strong><br>\nThis augmentation converts the image to the HLS color space, which separates the image into its hue, saturation, and brightness channels.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2HLS)\n</code></pre>\n<ul>\n<li><strong>LUV Color Space:</strong><br>\nThis augmentation converts the image to the LUV color space, which is designed to be perceptually uniform and enables more accurate color comparison.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2LUV)\n</code></pre>\n<ul>\n<li><strong>Alpha Channel:</strong><br>\nThis augmentation adds an alpha channel to the image, which can be used for transparency effects.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2RGBA)\n</code></pre>\n<ul>\n<li><strong>YZ Color Space:</strong><br>\nThis augmentation converts the image to the XYZ color space, which is a device-independent color space that allows for more accurate color representation.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2XYZ)\n</code></pre>\n<ul>\n<li><strong>Luma Chroma:</strong><br>\nThis augmentation converts the image to the YCrCb color space, which separates the image into its luma (brightness) and chroma (color) channels.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2YCrCb)\n</code></pre>\n<ul>\n<li><strong>CIE Lab:</strong></li>\n<li>This augmentation converts the image to the CIE Lab color space, which is designed to be perceptually uniform and enables more accurate color comparison.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2Lab)\n</code></pre>\n<ul>\n<li><strong>YUV Color Space:</strong><br>\nThis augmentation converts the image to the YUV color space, which separates the image into its luminance (brightness) and chrominance (color) channels.</li>\n</ul>\n<pre><code> ():\n     cv2.cvtColor(image, cv2.COLOR_RGB2YUV)\n</code></pre>\n<ul>\n<li><strong>Center Crop:</strong><br>\nThis augmentation randomly crops a rectangular area with an aspect ratio of [3/4,4/3], then randomly scales the crop by a factor between [8%,100%], and finally resizes the crop to an img_sizeXimg_size square. This is done randomly on each batch.</li>\n</ul>\n<pre><code>transforms.CenterCrop((, ))\n</code></pre>\n<p><strong>Flippings:</strong><br>\nThis augmentation adds a probability of random horizontal flip to the image. For example, with a probability of 0.5, the image has a 50% chance of being horizontally flipped.</p>\n<pre><code> ():\n     np.random.uniform() &lt; :\n        image = cv2.flip(image, )\n     image\n</code></pre>\n<ul>\n<li><strong>Random Crop:</strong><br>\nThis augmentation randomly crops a rectangular area from the image.</li>\n</ul>\n<pre><code>transforms.RandomCrop((, ))\n</code></pre>\n<ul>\n<li><strong>Random Resized Crop:</strong><br>\nThis augmentation randomly resizes and crops a rectangular area from the image.</li>\n</ul>\n<pre><code>transforms.RandomResizedCrop((, ))\n</code></pre>\n<ul>\n<li><strong>Color Jitter:</strong><br>\nThis augmentation randomly adjusts the brightness, contrast, saturation, and hue of the image.</li>\n</ul>\n<pre><code>transforms.ColorJitter(brightness=, contrast=, saturation=, hue=)\n</code></pre>\n<ul>\n<li><strong>Random Affine:</strong><br>\nThis augmentation randomly applies an affine transformation to the image, which includes rotation, scaling, and shearing.</li>\n</ul>\n<pre><code>transforms.RandomAffine(degrees=, translate=(, ), scale=(, ), shear=)\n</code></pre>\n<ul>\n<li><strong>Random Horizontal Flip:</strong><br>\nrandomly flips the image horizontally with a probability of 0.5.</li>\n</ul>\n<pre><code>transforms.RandomHorizontalFlip()\n</code></pre>\n<ul>\n<li><strong>Random Vertical Flip:</strong><br>\nThis augmentation randomly flips the image vertically with a probability of 0.5.</li>\n</ul>\n<pre><code>transforms.RandomVerticalFlip()\n</code></pre>\n<ul>\n<li><strong>Random Perspective:</strong><br>\nThis augmentation randomly applies a perspective transformation to the image.</li>\n</ul>\n<pre><code>transforms.RandomPerspective()\n</code></pre>\n<ul>\n<li><strong>Random Rotation:</strong><br>\nThis augmentation randomly rotates the image by a given degree range.</li>\n</ul>\n<pre><code>transforms.RandomRotation(degrees=)\n</code></pre>\n<ul>\n<li><strong>Random Invert:</strong><br>\nThis augmentation randomly inverts the colors of the image.</li>\n</ul>\n<pre><code>transforms.RandomInvert()\n</code></pre>\n<ul>\n<li><strong>Random Posterize:</strong><br>\nThis augmentation randomly reduces the number of bits used to represent each pixel value, resulting in a posterized effect.</li>\n</ul>\n<pre><code>transforms.RandomPosterize(bits=)\n</code></pre>\n<ul>\n<li><strong>Random Solarize:</strong><br>\nThis augmentation randomly applies a solarize effect to the image, where pixels above a certain intensity threshold are inverted.</li>\n</ul>\n<pre><code>transforms.RandomSolarize(threshold=)\n</code></pre>\n<ul>\n<li><strong>Random Autocontrast:</strong><br>\nThis augmentation randomly adjusts the contrast of the image by stretching the intensity values to the full available range.</li>\n</ul>\n<pre><code>transforms.RandomAutocontrast()\n</code></pre>\n<ul>\n<li><strong>Random Equalize:</strong><br>\nThis augmentation randomly equalizes the histogram of the image, resulting in increased contrast.</li>\n</ul>\n<pre><code>transforms.RandomEqualize()\n</code></pre>\n<h4>Advanced Augmentations</h4>\n<ul>\n<li><strong>Auto Augment:</strong><br>\nAuto Augment is an augmentation method that uses reinforcement learning to search for the optimal augmentation policies for a given dataset. It has been shown to improve the performance of image classification models.</li>\n</ul>\n<pre><code> autoaugment  AutoAugment\n\nauto_augment = AutoAugment()\nimage = auto_augment(image)\n</code></pre>\n<ul>\n<li><strong>Fast Autoaugment:</strong><br>\nFast Autoaugment is a faster implementation of the Auto Augment method. It uses a neural network to predict the optimal augmentation policies for a given dataset.</li>\n</ul>\n<pre><code> fast_autoaugment  FastAutoAugment\n\nfast_auto_augment = FastAutoAugment()\nimage = fast_auto_augment(image)\n</code></pre>\n<ul>\n<li><strong>Augmix:</strong><br>\nAugmix is an augmentation method that combines multiple augmented images to create a single, more diverse and realistic image. It has been shown to improve the robustness and generalization of image classification models.</li>\n</ul>\n<pre><code> augmix  AugMix\n\naug_mix = AugMix()\nimage = aug_mix(image)\n</code></pre>\n<ul>\n<li><strong>Mixup/Cutout:</strong><br>\nMixup is an augmentation method that combines two images by linearly interpolating their pixel values. Cutout is an augmentation method that randomly removes a rectangular region from the image. These methods have been shown to improve the robustness and generalization of image classification models.</li>\n</ul>\n<blockquote>\n  <ul>\n  <li>\"You take a picture of a cat and add some \"transparent dog\" on top of it. The amount of transparency is a hyperparam.\"</li>\n  </ul>\n</blockquote>\n<pre><code>x=*x1+(-)x2\ny=*x1+(-)y2\n</code></pre>\n<h3>Validation Time Augmentations</h3>\n<ul>\n<li>Adding validation set data augmentation will improve your model performance on test time augmentation and thus increase the absolution performance but if you seek to gain better performance use some of your validation set for data augmentation experiments or data augmentation experiments during training and use another part to increase your performance parameters, this will reduce your model overfitting.</li>\n<li>You can use the test time augmentation by running the same model multiple times.</li>\n<li>Run many validation data augmentation experiments at the same time to find the ones with the greatest impact your model performance.</li>\n<li><strong>USE EXPERIMENTS:</strong> Experiment. Experiment. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. </li>\n</ul>\n<h3>Test Time Augmentations</h3>\n<p>Augmentations can be useful not juse during training but also during test time. <br>\nSimply apply them on prediction and avg the results. </p>\n<ul>\n<li><strong>Predict on different img scales:</strong> Try to predict the same image but on different scales. </li>\n<li><strong>Color Standardization:</strong> Standardizatize as above.</li>\n</ul>\n<h3>Current High Scoring Models</h3>\n<p>Here is a summary of the current highest scoring models on this competition:</p>\n<p><strong>tf_efficientnetv2_s</strong><br>\nUsed on <a href=\"https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/theoviel\" target=\"_blank\">Theo Viel</a></p>\n<p><strong>seresnext50_32x4d</strong></p>\n<ul>\n<li>Find it on <a href=\"https://www.kaggle.com/code/vslaykovsky/infer-pytorch-aux-targets-weighted-loss-thres\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/vslaykovsky\" target=\"_blank\">Vladimir Slaykovskiy\n</a></li>\n<li>Also used on <a href=\"https://www.kaggle.com/code/hengck23/notebooke04a738685\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">Heng</a></li>\n</ul>\n<p><strong>tf_effv2_s_208_402</strong></p>\n<ul>\n<li>Find it on <a href=\"https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference\" target=\"_blank\">this</a> notebook by <a href=\"https://www.kaggle.com/radek1\" target=\"_blank\">Radek Osmulski</a></li>\n</ul>\n<h3>Models To Consider</h3>\n<p>And some ideas and models to try out.</p>\n<p><strong>Swin Transformer</strong></p>\n<ul>\n<li>Was the SOTA before EfficientNetV2, You can find example of it <a href=\"https://www.kaggle.com/phalanx/train-swin-t-pytorch-lightning/notebook\" target=\"_blank\">here</a></li>\n</ul>\n<p><strong>BeIT Transformer</strong></p>\n<ul>\n<li>Also known to be a very strong model. </li>\n</ul>\n<p><strong>ViT Transformers</strong></p>\n<ul>\n<li>This is the first vision transformer released. Still very powerful to this day.</li>\n</ul>\n<p><strong>EfficientNet L2</strong></p>\n<ul>\n<li>The largest efficientnet V1, If you can somehow get it to work in terms of compute power it might yield high performance.</li>\n</ul>\n<h3>Hidden Layers After Your Backbone</h3>\n<p>Adding more layers can be beneficial as you can use them to learn more advanced features, but it can also temper the fine-tuning of your large pretrained model.</p>\n<h3>Unfreeze Layer By Layer</h3>\n<ul>\n<li>A simple trick that can get you a tiny improvement would be to unfreeze layers of your pretrained backbone as the training progress. </li>\n<li><strong>Adding more layers and freezing everything else:</strong> As it turns out: Many solutions could be even further improved by incorporating another training phase after the pretrained model had been trained! This can be done by freezing the pretrained model and adding dense layers after it. You will get a tiny uplift in the upmost leaderboard with that change though. </li>\n</ul>\n<p><strong>Weight freezing and unfreezing in PyTorch</strong></p>\n<pre><code>\n param  model.parameters():\n  param.requires_grad =  \n</code></pre>\n<pre><code>\n param  model.parameters():\n  param.requires_grad =  \n</code></pre>\n<p><strong>Weight freezing and unfreezing in TensorFlow</strong></p>\n<pre><code>\nlayer.trainable = \n</code></pre>\n<pre><code>\nlayer.trainable = \n</code></pre>\n<h3>Learning-rates &amp; LR Shedulers</h3>\n<p>Learning rates and learning rate schedulers affect your models' training performance. Changing the learning rate can have a great impact on performance and also on training convergence. </p>\n<h4>Learning rate schedulers</h4>\n<p>In recent times, One Cycle Cosine schedule has shown to provide better results on multiple NLP tasks, you can use it this way:</p>\n<p><strong>One Cycle Cosine scheduling in PyTorch</strong></p>\n<pre><code> torch.optim.lr_scheduler  CosineAnnealingLR\noptimizer = torch.optim.AdamW(optimizer_grouped_parameters, lr=args.learning_rate, eps=args.adam_epsilon)\nscheduler = CosineAnnealingLR(optimizer, T_max=num_train_optimization_steps)\nnum_training_steps = num_train_optimization_steps / args.gradient_accumulation_steps\n</code></pre>\n<pre><code>\nscheduler.step()\n\n</code></pre>\n<p><strong>One Cycle Cosine scheduling in TensorFlow</strong></p>\n<pre><code>optimizer = tf.keras.optimizers.Adam(learning_rate)\nscheduler = tf.keras.optimizers.schedules.CosineDecay(learning_rate, decay_steps=num_training_steps)\n</code></pre>\n<h4>Some Learning Rate Schedulers Tips</h4>\n<ul>\n<li>Using “Triangular” or “One Cyclic” methods for learning rate scheduling can provide subtle but significant improvements - those intelligent methods of learning rate scheduling are can overcome some of the batch size issues and this will be obvious when using them, especially when plugging them into the multiple learning rate system. </li>\n<li>Take the time to research for the best learning rate scheduling method for your task and the models you are using, it is a very important part of how your model will converge. </li>\n<li>Learning rate schedules can be used to train models with lower batch sizes or multiple learning rates or any combination of those. </li>\n<li>Try low learning rates first to see if a dramatic learning rate increase will help or hurt your performance. </li>\n<li>Increasing the learning rate at later stages of training or multiple learning rates or high batch size or gradient accumulation or learning rate schedulers sometimes will help your model converge better, it is an advanced technique as sometimes it can hurt the performance but only if you give it too large a value - remember to test it. </li>\n<li>Loss scaling can help reduce loss variance and improve gradient flow when using gradient accumulation or multiple learning rates or high batch size, but if you are trying to solve that problem by increasing batch size, try to increase the learning rate instead as it can sometimes yield better performance. </li>\n<li>In recent times, one cycle learning rate scheduling approaches have taken the leading positions on public leaderboards and past winning solutions. </li>\n</ul>\n<h4>Hyperparameters Of Optimizers</h4>\n<blockquote>\n  <p>Crossposting from \"Transformers Hyperparams Tips\"</p>\n</blockquote>\n<p>There are several things you need to know if you seek to get the best performance out of Adam optimizer:</p>\n<ul>\n<li>Finding the best weight decay value can be tricky and experiments (and luck) will be your best friends here. </li>\n<li>Another important hyperparameter is the <code>beta1</code> and <code>beta2</code> used in the Adam optimizer, choosing the best values depends on your task and the data. many new tasks can benefit from lower beta1 and higher beta2 while on established tasks they would perform the opposite. So again: Experiments will be your best friend here.</li>\n<li>In the world of Adam optimizer, the number one rule is not to underestimate the importance of the optimizer epsilon value. The same principle of finding the best weight decay hyperparameter applies here.</li>\n<li>Do not overuse the gradient clipping norm - it might sometimes be helpful when your gradients explode while the opposite is also true - it can prevent convergence on some tasks.</li>\n<li>Gradient accumulation still sows to provide some subtle benefits, I usually accumulate gradients for about 2 steps but if your GPU does not run out of memory, you can push up to 8 gradient accumulation steps. Gradient accumulation is also useful when using mixed-precision. </li>\n</ul>\n<h3>Overfitting and regularization</h3>\n<ul>\n<li>Use dropout! adding layers dropouts between layers usually yields more training stability and more robust results, use dropouts with your hidden layers.</li>\n<li>Dropout can also be used to increase performance by small margins, experiment with setting layer dropouts before training. the task and the model.</li>\n<li>**If you are into regularization: **Regularization can provide marge uplift on performance when your NN is overfitting or underfitting, for normal Machine learning models, L1 or L2 regularization are OK but also try to use additive and hidden layers dropout.</li>\n<li><strong>ALWAYS TEST OUT IDEAS USING EXPERIMENTS:</strong> Use experiments. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. </li>\n<li><strong>Multi Validations:</strong> You can increase your model's robustness to overfitting by using multiple validations. This however comes at the cost of compute time.</li>\n</ul>\n<h3>Optimizer</h3>\n<p>Everyone today are using Adam or AdamW. BUT something to consider is often times if you hyperparam-search SGD with momentum enough, you probably could get better results with it. But again this requires heavy tuning.</p>\n<p><strong>There are several notable optimizers that are important to know about:</strong></p>\n<ul>\n<li><strong>AdamW:</strong>  This is an extension to the Adam algorithm that prevents exponential weights decay of the model's weights in the outer layers as well as encourages penalized hyper volume below the default weights. </li>\n<li><strong>Adafactor:</strong> It was designed to have low memory usage and is scalable. This optimizer can offer significant optimizer performances using multiple GPUs (see below). </li>\n<li><strong>Novograd:</strong> Basically another Adam-like optimizer but with very better properties. It is one of the optimizers used by  to train the bert-large model. </li>\n<li><strong>Ranger:</strong> Ranger optimizer is a very interesting optimizer that has pretty good results in winning solutions it terms of performance optimization but it is obviously not very well known or supported so I leave that up to you to research. </li>\n<li><strong>Lamb:</strong> GPU optimized reusable Adam optimizer developed by the GLUE and QQP competition winner. </li>\n<li><strong>Lookahead:</strong> A popular optimizer you can use on top of other optimizers and it will provide you with some performance gains. </li>\n</ul>\n<h3>Label Smoothing</h3>\n<blockquote>\n  <p>Original Paper: <a href=\"https://arxiv.org/pdf/1906.02629.pdf\" target=\"_blank\">https://arxiv.org/pdf/1906.02629.pdf</a></p>\n</blockquote>\n<p>Basiclly a simple trick: y_true = y_true * (1.0 - ε) + 0.5 * ε [example: ε = 0.001]<br>\nUsually works well. You can find it on some of the high scoring public notebook of this competition.</p>\n<p><strong>Tensorflow:</strong> </p>\n<pre><code>loss = BinaryCrossentropy(label_smoothing = label_smoothing)\n</code></pre>\n<p><strong>Pytorch:</strong></p>\n<pre><code> torch.nn.modules.loss  _WeightedLoss\n\n ():\n     ():\n        ().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n        self.pos_weight = pos_weight\n\n\n     ():\n          &lt;= smoothing &lt; \n         torch.no_grad(): targets = targets * ( - smoothing) +  * smoothing\n         targets\n\n     ():\n        targets = SmoothBCEwLogits._smooth(targets, inputs.size(-), self.smoothing)\n        loss = F.binary_cross_entropy_with_logits(inputs, targets,self.weight, pos_weight = self.pos_weight)\n          self.reduction == : loss = loss.()\n          self.reduction == : loss = loss.mean()\n         loss\n</code></pre>\n<pre><code>    \n</code></pre>\n<h2>Advanced Tricks</h2>\n<h3>Knowledge Distillation</h3>\n<blockquote>\n  <p>Use a large teacher network to guide the learning of a small network.</p>\n</blockquote>\n<ul>\n<li><strong>Train the large model:</strong> Train a large model on the data.</li>\n<li><strong>Calculate soft target:</strong> Use the trained large model to calculate soft target. That is, the output of the softmax after the large model is \"softened\"</li>\n<li><strong>Studnet Model Training:</strong> Train a student model based on the teachers output as an extra soft target loss function on the basis of the large model, and adjust the proportion of the two loss functions through interpulation.</li>\n</ul>\n<h3>Pseudo Labeling</h3>\n<blockquote>\n  <p>Use a model to label unlabeled data (For example test data) then use the new labeled data you have for training your models.</p>\n</blockquote>\n<ul>\n<li><strong>Train the teacher model:</strong> Train a model on the data you have.</li>\n<li><strong>Calculate soft target:</strong> Use the trained large model to calculate soft target for unlabeled data.</li>\n<li><strong>Important:</strong> Use only the targets your model is \"sure\" about** Use only the highest confidence of the OOF predictions as pseudo labels to avoid mistakes as much as possible. (It might not work if you don't do this.).</li>\n<li><strong>Studnet Model Training:</strong> Train a student model on the new labeled data that you have.</li>\n</ul>\n<blockquote>\n  <p><strong>How do you CV pseudo labeling?</strong> Simple: You concat the unlabeled data to your OOF and simply use them just as you would have used stacking. </p>\n</blockquote>\n<h3>Error Analysis</h3>\n<p>An important practice that can save you a lot of time is to use your model for finding harder or broken data samples.<br>\nThere can be many reasons for an image to be \"harder\" for your model, for example, small target objects, different coloring, cut off target, invalid annotations and more.</p>\n<p><strong>Mistakes are Good News!</strong></p>\n<p>These are exactly the samples that separate the top of the leaderboard from the rest of the participants.<br>\nIf you have a hard time explaining what is going on with your model, it might be a good idea to take a look at the validation samples your model struggles with.</p>\n<h3>Finding Your Model's Errors</h3>\n<p>The easiest way to find errors is to sort validation samples by the model's confidence score and see which ones are being predicted with the lowest confidence.</p>\n<pre><code>mistakes_idx = [img_idx  img_idx  ((train))  (pred[img_idx] &gt; ) != target[img_idx]]\nmistakes_preds = pred[mistakes_idx]\nsorted_idx = np.argsort(mistakes_preds)[:]\n\n</code></pre>",
      "rawMarkdown": "The following are image classification tips and tricks you might find useful for this competition. \n\nEnjoy! \n_____\n\n## Image Classification Tips & Tricks\n\n#### Augmentations\n\n#### Color Augmentations\n\n- **Color Skew:**\nThis augmentation randomly adjusts the hue, saturation, and brightness of the image by multiplying each channel by a randomly chosen coefficient. The coefficients are chosen from a range of [0:6;1:4] to ensure that the resulting image is not too distorted.\n\n```python\ndef color_skew(image):\n    h, s, v = cv2.split(image)\n    h = h * np.random.uniform(low=0, high=6)\n    s = s * np.random.uniform(low=1, high=4)\n    v = v * np.random.uniform(low=0, high=6)\n    return cv2.merge((h, s, v))\n```\n\n- **RGB Norm:**\nThis augmentation normalizes the RGB channels of the image by subtracting the mean value of each channel from the values in that channel and dividing by the standard deviation of the channel. This helps to standardize the values in the image and can improve the performance of the model.\n\n```python\ndef rgb_norm(image):\n    r, g, b = cv2.split(image)\n    r = (r - np.mean(r)) / np.std(r)\n    g = (g - np.mean(g)) / np.std(g)\n    b = (b - np.mean(b)) / np.std(b)\n    return cv2.merge((r, g, b))\n```\n\n- **Black and White:**\nThis augmentation converts the image to black and white by converting it to the grayscale color space.\n\n```python\ndef black_and_white(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n```\n\n- **Ben Graham: Greyscale + Gaussian Blur:**\nThis augmentation converts the image to grayscale and applies a Gaussian blur to smooth out any noise or details in the image.\n\n```python\ndef ben_graham(image):\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n    image = cv2.GaussianBlur(image, (5, 5), 0)\n    return image\n```\n\n- **Hue, Saturation, Brightness:**\nThis augmentation converts the image to the HLS color space, which separates the image into its hue, saturation, and brightness channels.\n\n```python\ndef hsb(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2HLS)\n```\n\n- **LUV Color Space:**\nThis augmentation converts the image to the LUV color space, which is designed to be perceptually uniform and enables more accurate color comparison.\n\n```python\ndef luv(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2LUV)\n```\n\n- **Alpha Channel:**\nThis augmentation adds an alpha channel to the image, which can be used for transparency effects.\n\n```python\ndef alpha_channel(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2RGBA)\n```\n\t\n- **YZ Color Space:**\nThis augmentation converts the image to the XYZ color space, which is a device-independent color space that allows for more accurate color representation.\n\n```python\ndef xyz(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2XYZ)\n```\n\n- **Luma Chroma:**\nThis augmentation converts the image to the YCrCb color space, which separates the image into its luma (brightness) and chroma (color) channels.\n\n```python\ndef luma_chroma(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2YCrCb)\n```\n\n- **CIE Lab:**\n- This augmentation converts the image to the CIE Lab color space, which is designed to be perceptually uniform and enables more accurate color comparison.\n\n```python\ndef cie_lab(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2Lab)\n```\n\n- **YUV Color Space:**\nThis augmentation converts the image to the YUV color space, which separates the image into its luminance (brightness) and chrominance (color) channels.\n\n```python\ndef yuv(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2YUV)\n```\n\n- **Center Crop:**\nThis augmentation randomly crops a rectangular area with an aspect ratio of [3/4,4/3], then randomly scales the crop by a factor between [8%,100%], and finally resizes the crop to an img_sizeXimg_size square. This is done randomly on each batch.\n\n```python\ntransforms.CenterCrop((100, 100))\n```\n\n**Flippings:**\nThis augmentation adds a probability of random horizontal flip to the image. For example, with a probability of 0.5, the image has a 50% chance of being horizontally flipped.\n\n```python\ndef flippings(image):\n    if np.random.uniform() < 0.5:\n        image = cv2.flip(image, 1)\n    return image\n```\n\n- **Random Crop:**\n\tThis augmentation randomly crops a rectangular area from the image.\n\n```python\ntransforms.RandomCrop((100, 100))\n```\n\n- **Random Resized Crop:**\nThis augmentation randomly resizes and crops a rectangular area from the image.\n\n```python\ntransforms.RandomResizedCrop((100, 100))\n```\n\n- **Color Jitter:**\nThis augmentation randomly adjusts the brightness, contrast, saturation, and hue of the image.\n\n```python\ntransforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5)\n```\n\n- **Random Affine:**\nThis augmentation randomly applies an affine transformation to the image, which includes rotation, scaling, and shearing.\n\n```python\ntransforms.RandomAffine(degrees=45, translate=(0.1, 0.1), scale=(0.5, 2.0), shear=45)\n```\n\n- **Random Horizontal Flip:**\nrandomly flips the image horizontally with a probability of 0.5.\n\n```python\ntransforms.RandomHorizontalFlip()\n```\n\n- **Random Vertical Flip:**\nThis augmentation randomly flips the image vertically with a probability of 0.5.\n\n```python\ntransforms.RandomVerticalFlip()\n```\n\n- **Random Perspective:**\nThis augmentation randomly applies a perspective transformation to the image.\n\n```python\ntransforms.RandomPerspective()\n```\n\n- **Random Rotation:**\nThis augmentation randomly rotates the image by a given degree range.\n\n```python\ntransforms.RandomRotation(degrees=45)\n```\n\n- **Random Invert:**\nThis augmentation randomly inverts the colors of the image.\n\n```python\ntransforms.RandomInvert()\n```\n\n- **Random Posterize:**\nThis augmentation randomly reduces the number of bits used to represent each pixel value, resulting in a posterized effect.\n\n```python\ntransforms.RandomPosterize(bits=4)\n```\n\n- **Random Solarize:**\nThis augmentation randomly applies a solarize effect to the image, where pixels above a certain intensity threshold are inverted.\n\n```python\ntransforms.RandomSolarize(threshold=128)\n```\n\n- **Random Autocontrast:**\nThis augmentation randomly adjusts the contrast of the image by stretching the intensity values to the full available range.\n\n```python\ntransforms.RandomAutocontrast()\n```\n\n- **Random Equalize:**\nThis augmentation randomly equalizes the histogram of the image, resulting in increased contrast.\n\n```python\ntransforms.RandomEqualize()\n```\n\n#### Advanced Augmentations\n\n- **Auto Augment:**\nAuto Augment is an augmentation method that uses reinforcement learning to search for the optimal augmentation policies for a given dataset. It has been shown to improve the performance of image classification models.\n\n```python\nfrom autoaugment import AutoAugment\n\nauto_augment = AutoAugment()\nimage = auto_augment(image)\n```\n\n- **Fast Autoaugment:**\nFast Autoaugment is a faster implementation of the Auto Augment method. It uses a neural network to predict the optimal augmentation policies for a given dataset.\n\n```python\nfrom fast_autoaugment import FastAutoAugment\n\nfast_auto_augment = FastAutoAugment()\nimage = fast_auto_augment(image)\n```\n\n- **Augmix:**\nAugmix is an augmentation method that combines multiple augmented images to create a single, more diverse and realistic image. It has been shown to improve the robustness and generalization of image classification models.\n\n```python\nfrom augmix import AugMix\n\naug_mix = AugMix()\nimage = aug_mix(image)\n```\n\n- **Mixup/Cutout:**\nMixup is an augmentation method that combines two images by linearly interpolating their pixel values. Cutout is an augmentation method that randomly removes a rectangular region from the image. These methods have been shown to improve the robustness and generalization of image classification models.\n\n> - \"You take a picture of a cat and add some \"transparent dog\" on top of it. The amount of transparency is a hyperparam.\"\n\n```python\nx=lambda*x1+(1-lambda)x2\ny=lambda*x1+(1-lambda)y2\n```\n\n### Validation Time Augmentations\n\n* Adding validation set data augmentation will improve your model performance on test time augmentation and thus increase the absolution performance but if you seek to gain better performance use some of your validation set for data augmentation experiments or data augmentation experiments during training and use another part to increase your performance parameters, this will reduce your model overfitting.\n* You can use the test time augmentation by running the same model multiple times.\n* Run many validation data augmentation experiments at the same time to find the ones with the greatest impact your model performance.\n* **USE EXPERIMENTS:** Experiment. Experiment. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. \n\n### Test Time Augmentations\nAugmentations can be useful not juse during training but also during test time. \nSimply apply them on prediction and avg the results. \n\n- **Predict on different img scales:** Try to predict the same image but on different scales. \n- **Color Standardization:** Standardizatize as above.\n\n\n### Current High Scoring Models\n\nHere is a summary of the current highest scoring models on this competition:\n\n**tf_efficientnetv2_s**\nUsed on [this](https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali) notebook by [Theo Viel](https://www.kaggle.com/theoviel)\n\n**seresnext50_32x4d**\n- Find it on [this](https://www.kaggle.com/code/vslaykovsky/infer-pytorch-aux-targets-weighted-loss-thres) notebook by [Vladimir Slaykovskiy\n](https://www.kaggle.com/vslaykovsky)\n- Also used on [this](https://www.kaggle.com/code/hengck23/notebooke04a738685) notebook by [Heng](https://www.kaggle.com/hengck23)\n\n**tf_effv2_s_208_402**\n- Find it on [this](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference) notebook by [Radek Osmulski](https://www.kaggle.com/radek1)\n\n\n### Models To Consider\n\nAnd some ideas and models to try out.\n\n**Swin Transformer**\n\n- Was the SOTA before EfficientNetV2, You can find example of it [here](https://www.kaggle.com/phalanx/train-swin-t-pytorch-lightning/notebook)\n\n**BeIT Transformer**\n\n- Also known to be a very strong model. \n\n**ViT Transformers**\n\n- This is the first vision transformer released. Still very powerful to this day.\n\n**EfficientNet L2**\n\n- The largest efficientnet V1, If you can somehow get it to work in terms of compute power it might yield high performance.\n\n### Hidden Layers After Your Backbone\n\nAdding more layers can be beneficial as you can use them to learn more advanced features, but it can also temper the fine-tuning of your large pretrained model.\n\n### Unfreeze Layer By Layer\n- A simple trick that can get you a tiny improvement would be to unfreeze layers of your pretrained backbone as the training progress. \n- **Adding more layers and freezing everything else:** As it turns out: Many solutions could be even further improved by incorporating another training phase after the pretrained model had been trained! This can be done by freezing the pretrained model and adding dense layers after it. You will get a tiny uplift in the upmost leaderboard with that change though. \n\n**Weight freezing and unfreezing in PyTorch**\n\n```python\n# Weight freezing\nfor param in model.parameters():\n  param.requires_grad = False # unfreeze weights in at all\n```\n\n```python\n# Weight unfreezing\nfor param in model.parameters():\n  param.requires_grad = True # unfreeze weights in at all\n```\n\n**Weight freezing and unfreezing in TensorFlow**\n\n```python\n# Weight freezing\nlayer.trainable = False\n```\n```python\n# Weight unfreezing\nlayer.trainable = True\n```\n\n### Learning-rates & LR Shedulers\n\nLearning rates and learning rate schedulers affect your models' training performance. Changing the learning rate can have a great impact on performance and also on training convergence. \n\n#### Learning rate schedulers\n\nIn recent times, One Cycle Cosine schedule has shown to provide better results on multiple NLP tasks, you can use it this way:\n\n**One Cycle Cosine scheduling in PyTorch**\n\n```python\nfrom torch.optim.lr_scheduler import CosineAnnealingLR\noptimizer = torch.optim.AdamW(optimizer_grouped_parameters, lr=args.learning_rate, eps=args.adam_epsilon)\nscheduler = CosineAnnealingLR(optimizer, T_max=num_train_optimization_steps)\nnum_training_steps = num_train_optimization_steps / args.gradient_accumulation_steps\n```\n\n```python\n# Update the scheduler\nscheduler.step()\n# step the learning rate scheduler here, you will want to step the learning rate scheduler only once per optimizer step nothing more nothing less. So in this case, it should be called before you expect the gradients to be applied.\n```\n\n**One Cycle Cosine scheduling in TensorFlow**\n\n```python\noptimizer = tf.keras.optimizers.Adam(learning_rate)\nscheduler = tf.keras.optimizers.schedules.CosineDecay(learning_rate, decay_steps=num_training_steps)\n```\n\n#### Some Learning Rate Schedulers Tips\n\n* Using “Triangular” or “One Cyclic” methods for learning rate scheduling can provide subtle but significant improvements - those intelligent methods of learning rate scheduling are can overcome some of the batch size issues and this will be obvious when using them, especially when plugging them into the multiple learning rate system. \n* Take the time to research for the best learning rate scheduling method for your task and the models you are using, it is a very important part of how your model will converge. \n* Learning rate schedules can be used to train models with lower batch sizes or multiple learning rates or any combination of those. \n* Try low learning rates first to see if a dramatic learning rate increase will help or hurt your performance. \n* Increasing the learning rate at later stages of training or multiple learning rates or high batch size or gradient accumulation or learning rate schedulers sometimes will help your model converge better, it is an advanced technique as sometimes it can hurt the performance but only if you give it too large a value - remember to test it. \n* Loss scaling can help reduce loss variance and improve gradient flow when using gradient accumulation or multiple learning rates or high batch size, but if you are trying to solve that problem by increasing batch size, try to increase the learning rate instead as it can sometimes yield better performance. \n* In recent times, one cycle learning rate scheduling approaches have taken the leading positions on public leaderboards and past winning solutions. \n\n#### Hyperparameters Of Optimizers\n\n> Crossposting from \"Transformers Hyperparams Tips\"\n\nThere are several things you need to know if you seek to get the best performance out of Adam optimizer:\n\n* Finding the best weight decay value can be tricky and experiments (and luck) will be your best friends here. \n* Another important hyperparameter is the `beta1` and `beta2` used in the Adam optimizer, choosing the best values depends on your task and the data. many new tasks can benefit from lower beta1 and higher beta2 while on established tasks they would perform the opposite. So again: Experiments will be your best friend here.\n* In the world of Adam optimizer, the number one rule is not to underestimate the importance of the optimizer epsilon value. The same principle of finding the best weight decay hyperparameter applies here.\n* Do not overuse the gradient clipping norm - it might sometimes be helpful when your gradients explode while the opposite is also true - it can prevent convergence on some tasks.\n* Gradient accumulation still sows to provide some subtle benefits, I usually accumulate gradients for about 2 steps but if your GPU does not run out of memory, you can push up to 8 gradient accumulation steps. Gradient accumulation is also useful when using mixed-precision. \n\n### Overfitting and regularization\n\n* Use dropout! adding layers dropouts between layers usually yields more training stability and more robust results, use dropouts with your hidden layers.\n* Dropout can also be used to increase performance by small margins, experiment with setting layer dropouts before training. the task and the model.\n* **If you are into regularization: **Regularization can provide marge uplift on performance when your NN is overfitting or underfitting, for normal Machine learning models, L1 or L2 regularization are OK but also try to use additive and hidden layers dropout.\n* **ALWAYS TEST OUT IDEAS USING EXPERIMENTS:** Use experiments. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. \n* **Multi Validations:** You can increase your model's robustness to overfitting by using multiple validations. This however comes at the cost of compute time.\n\n### Optimizer\nEveryone today are using Adam or AdamW. BUT something to consider is often times if you hyperparam-search SGD with momentum enough, you probably could get better results with it. But again this requires heavy tuning.\n\n**There are several notable optimizers that are important to know about:**\n- **AdamW:**  This is an extension to the Adam algorithm that prevents exponential weights decay of the model's weights in the outer layers as well as encourages penalized hyper volume below the default weights. \n- **Adafactor:** It was designed to have low memory usage and is scalable. This optimizer can offer significant optimizer performances using multiple GPUs (see below). \n- **Novograd:** Basically another Adam-like optimizer but with very better properties. It is one of the optimizers used by  to train the bert-large model. \n- **Ranger:** Ranger optimizer is a very interesting optimizer that has pretty good results in winning solutions it terms of performance optimization but it is obviously not very well known or supported so I leave that up to you to research. \n- **Lamb:** GPU optimized reusable Adam optimizer developed by the GLUE and QQP competition winner. \n- **Lookahead:** A popular optimizer you can use on top of other optimizers and it will provide you with some performance gains. \n\n### Label Smoothing\n\n> Original Paper: https://arxiv.org/pdf/1906.02629.pdf\n\nBasiclly a simple trick: y_true = y_true * (1.0 - ε) + 0.5 * ε [example: ε = 0.001]\nUsually works well. You can find it on some of the high scoring public notebook of this competition.\n\n\n**Tensorflow:** \n```python\nloss = BinaryCrossentropy(label_smoothing = label_smoothing)\n```\n\n**Pytorch:**\n\n```python\nfrom torch.nn.modules.loss import _WeightedLoss\n\nclass SmoothBCEwLogits(_WeightedLoss):\n    def __init__(self, weight = None, reduction = 'mean', smoothing = 0.0, pos_weight = None):\n        super().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n        self.pos_weight = pos_weight\n\n    @staticmethod\n    def _smooth(targets, n_labels, smoothing = 0.0):\n        assert 0 <= smoothing < 1\n        with torch.no_grad(): targets = targets * (1.0 - smoothing) + 0.5 * smoothing\n        return targets\n\n    def forward(self, inputs, targets):\n        targets = SmoothBCEwLogits._smooth(targets, inputs.size(-1), self.smoothing)\n        loss = F.binary_cross_entropy_with_logits(inputs, targets,self.weight, pos_weight = self.pos_weight)\n        if  self.reduction == 'sum': loss = loss.sum()\n        elif  self.reduction == 'mean': loss = loss.mean()\n        return loss\n```\t\t\n\n\n\n\n## Advanced Tricks\n\n### Knowledge Distillation\n> Use a large teacher network to guide the learning of a small network.\n\n- **Train the large model:** Train a large model on the data.\n- **Calculate soft target:** Use the trained large model to calculate soft target. That is, the output of the softmax after the large model is \"softened\"\n- **Studnet Model Training:** Train a student model based on the teachers output as an extra soft target loss function on the basis of the large model, and adjust the proportion of the two loss functions through interpulation.\n\n\n### Pseudo Labeling\n> Use a model to label unlabeled data (For example test data) then use the new labeled data you have for training your models.\n\n- **Train the teacher model:** Train a model on the data you have.\n- **Calculate soft target:** Use the trained large model to calculate soft target for unlabeled data.\n- **Important:** Use only the targets your model is \"sure\" about** Use only the highest confidence of the OOF predictions as pseudo labels to avoid mistakes as much as possible. (It might not work if you don't do this.).\n- **Studnet Model Training:** Train a student model on the new labeled data that you have.\n\n> **How do you CV pseudo labeling?** Simple: You concat the unlabeled data to your OOF and simply use them just as you would have used stacking. \n\n\n### Error Analysis\n\nAn important practice that can save you a lot of time is to use your model for finding harder or broken data samples.\nThere can be many reasons for an image to be \"harder\" for your model, for example, small target objects, different coloring, cut off target, invalid annotations and more.\n\n**Mistakes are Good News!**\n\nThese are exactly the samples that separate the top of the leaderboard from the rest of the participants.\nIf you have a hard time explaining what is going on with your model, it might be a good idea to take a look at the validation samples your model struggles with.\n\n### Finding Your Model's Errors\n\nThe easiest way to find errors is to sort validation samples by the model's confidence score and see which ones are being predicted with the lowest confidence.\n\n```python\nmistakes_idx = [img_idx for img_idx in range(len(train)) if int(pred[img_idx] > 0.5) != target[img_idx]]\nmistakes_preds = pred[mistakes_idx]\nsorted_idx = np.argsort(mistakes_preds)[:20]\n# Show the images of the sorted idx here..\n```\n\n\n",
      "votes": 188
    },
    {
      "id": 2067445,
      "postDate": "2022-12-16T17:56:04.983Z",
      "content": "<p>Of course, the above suggestion would be really helpful. But I've seen many similar post on each image competition. By doing so, it becomes easy for user to go from novice - grad-master tier in discussion. <strong>Kaggle should place these common suggestion to their learning section</strong> <a href=\"https://www.kaggle.com/learn/computer-vision\" target=\"_blank\">https://www.kaggle.com/learn/computer-vision</a>.</p>\n<p><strong>If the tips and tricks are more oriented to a competition, only then it is purely useful.</strong> For example, tips and tricks regarding of how to tackle imbalance dataset, or bigger input size training etc with limited resource - which are actual issues in this competition. And that also should be provided with some results and not some text from google. </p>",
      "rawMarkdown": "Of course, the above suggestion would be really helpful. But I've seen many similar post on each image competition. By doing so, it becomes easy for user to go from novice - grad-master tier in discussion. **Kaggle should place these common suggestion to their learning section** https://www.kaggle.com/learn/computer-vision.\n\n**If the tips and tricks are more oriented to a competition, only then it is purely useful.** For example, tips and tricks regarding of how to tackle imbalance dataset, or bigger input size training etc with limited resource - which are actual issues in this competition. And that also should be provided with some results and not some text from google. ",
      "votes": 6,
      "replies": [
        {
          "id": 2067633,
          "postDate": "2022-12-16T23:24:52.887Z",
          "content": "<p><a href=\"https://www.kaggle.com/simonalerdic\" target=\"_blank\">@simonalerdic</a> </p>\n<p>Honest question, can you provide some links to the 'many similar posts' you've seen?   I'd like to add them to my bookmark list.</p>\n<p>That said, I agree that the author should also post a copy of this in the learn/computer-vision section.</p>",
          "rawMarkdown": "@simonalerdic \n\nHonest question, can you provide some links to the 'many similar posts' you've seen?   I'd like to add them to my bookmark list.\n\nThat said, I agree that the author should also post a copy of this in the learn/computer-vision section.",
          "votes": 1
        },
        {
          "id": 2068523,
          "postDate": "2022-12-18T04:42:52.230Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2070155,
              "postDate": "2022-12-19T17:03:35.323Z",
              "content": "<p>me agree too</p>",
              "rawMarkdown": "me agree too"
            }
          ]
        }
      ]
    },
    {
      "id": 2068137,
      "postDate": "2022-12-17T14:55:31.883Z",
      "content": "<p>Awesome tutorial 👍</p>\n<p>I think there's a little mistake in the ben_graham function.</p>\n<p>The image is not in grayscale here. You can therefore replace the color transform function or select only the value before blurring.</p>\n<pre><code> ():\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n  image = cv2.GaussianBlur(image, (,), )\n   image\n</code></pre>\n<p>or</p>\n<pre><code> ():\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n  image = cv2.GaussianBlur(image[], (,), )\n   image\n</code></pre>",
      "rawMarkdown": "Awesome tutorial 👍\n\nI think there's a little mistake in the ben_graham function.\n\nThe image is not in grayscale here. You can therefore replace the color transform function or select only the value before blurring.\n\n```python\ndef ben_graham(image):\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n  image = cv2.GaussianBlur(image, (5,5), 0)\n  return image\n```\nor\n\n```python\ndef ben_graham(image):\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n  image = cv2.GaussianBlur(image[2], (5,5), 0)\n  return image\n```",
      "votes": 2,
      "replies": [
        {
          "id": 2068525,
          "postDate": "2022-12-18T04:43:46.937Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2327695,
      "postDate": "2023-07-03T05:47:15.520Z",
      "content": "<p>Good Job ~!</p>\n<p>Realtime Histogram of Lab Colorspace is here: <a href=\"https://youtu.be/DXHI5xv6FF8\" target=\"_blank\">https://youtu.be/DXHI5xv6FF8</a></p>",
      "rawMarkdown": "Good Job ~!\n\nRealtime Histogram of Lab Colorspace is here: https://youtu.be/DXHI5xv6FF8"
    },
    {
      "id": 2148812,
      "postDate": "2023-02-17T17:28:38.823Z",
      "content": "<p>This is incredibly useful. Thank you so much!!</p>",
      "rawMarkdown": "This is incredibly useful. Thank you so much!!"
    },
    {
      "id": 2134695,
      "postDate": "2023-02-08T07:44:46.503Z",
      "content": "<p>Thanks for your sharing. <br>\nI wonder whether the 'Augmentor' library could be used in this competition.<br>\nHave you tried that?</p>",
      "rawMarkdown": "Thanks for your sharing. \nI wonder whether the 'Augmentor' library could be used in this competition.\nHave you tried that?"
    },
    {
      "id": 2085950,
      "postDate": "2023-01-04T13:34:05.903Z",
      "content": "<p>A very nice tutorial thanks a lot for sharing this amazing material with us.</p>",
      "rawMarkdown": "A very nice tutorial thanks a lot for sharing this amazing material with us."
    },
    {
      "id": 2070018,
      "postDate": "2022-12-19T14:35:11.963Z",
      "content": "<p>The presentation is great</p>",
      "rawMarkdown": "The presentation is great"
    },
    {
      "id": 2069804,
      "postDate": "2022-12-19T10:40:57.070Z",
      "content": "<p>Bookmarking this, really nice work <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>!</p>",
      "rawMarkdown": "Bookmarking this, really nice work @thedevastator!"
    },
    {
      "id": 2067634,
      "postDate": "2022-12-16T23:27:06.043Z",
      "content": "<p>Thanks, the post was very devastating!  :)</p>",
      "rawMarkdown": "Thanks, the post was very devastating!  :)"
    },
    {
      "id": 2067401,
      "postDate": "2022-12-16T17:06:38.960Z",
      "content": "<p>Outstanding! If I could add more votes it would be at least 50 :) <br>\nBTW: nice avatar :) minimalistic. </p>",
      "rawMarkdown": "Outstanding! If I could add more votes it would be at least 50 :) \nBTW: nice avatar :) minimalistic. "
    },
    {
      "id": 2067421,
      "postDate": "2022-12-16T17:21:07.233Z",
      "content": "<p>Very nice thanks for sharing. </p>",
      "rawMarkdown": "Very nice thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 2133510,
      "postDate": "2023-02-07T13:07:51.523Z",
      "content": "<p>Very useful content! Thank you so much.</p>",
      "rawMarkdown": "Very useful content! Thank you so much."
    },
    {
      "id": 2071353,
      "postDate": "2022-12-20T23:32:57.667Z",
      "content": "<p>It's useful, Thank you for your effort.</p>",
      "rawMarkdown": "It's useful, Thank you for your effort."
    },
    {
      "id": 2071146,
      "postDate": "2022-12-20T17:29:13.733Z",
      "content": "<p>Thanks. This was helpful.</p>",
      "rawMarkdown": "Thanks. This was helpful."
    },
    {
      "id": 2070770,
      "postDate": "2022-12-20T10:01:52.680Z",
      "content": "<p>This is helpful, thanks for sharing</p>",
      "rawMarkdown": "This is helpful, thanks for sharing"
    },
    {
      "id": 2070587,
      "postDate": "2022-12-20T07:18:26.093Z",
      "content": "<p>Thanks fot the tips</p>",
      "rawMarkdown": "Thanks fot the tips"
    },
    {
      "id": 2068033,
      "postDate": "2022-12-17T12:47:39.033Z",
      "content": "<p>Very useful !!! Thanks for sharing</p>",
      "rawMarkdown": "Very useful !!! Thanks for sharing"
    },
    {
      "id": 2067629,
      "postDate": "2022-12-16T23:11:56.293Z",
      "content": "<p>Awesome, thanks.</p>",
      "rawMarkdown": "Awesome, thanks."
    }
  ],
  "comments": [
    {
      "id": 2067445,
      "author_name": "Simon Alerdic",
      "author_url": "",
      "post_date": "2022-12-16T17:56:04.983000",
      "content": "<p>Of course, the above suggestion would be really helpful. But I've seen many similar post on each image competition. By doing so, it becomes easy for user to go from novice - grad-master tier in discussion. <strong>Kaggle should place these common suggestion to their learning section</strong> <a href=\"https://www.kaggle.com/learn/computer-vision\" target=\"_blank\">https://www.kaggle.com/learn/computer-vision</a>.</p>\n<p><strong>If the tips and tricks are more oriented to a competition, only then it is purely useful.</strong> For example, tips and tricks regarding of how to tackle imbalance dataset, or bigger input size training etc with limited resource - which are actual issues in this competition. And that also should be provided with some results and not some text from google. </p>",
      "votes": 6,
      "replies": [
        {
          "id": 2067633,
          "author_name": "@kaggleqrdl",
          "author_url": "",
          "post_date": "2022-12-16T23:24:52.887000",
          "content": "<p><a href=\"https://www.kaggle.com/simonalerdic\" target=\"_blank\">@simonalerdic</a> </p>\n<p>Honest question, can you provide some links to the 'many similar posts' you've seen?   I'd like to add them to my bookmark list.</p>\n<p>That said, I agree that the author should also post a copy of this in the learn/computer-vision section.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2068523,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-18T04:42:52.230000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2070155,
              "author_name": "TrappedInTessaract",
              "author_url": "",
              "post_date": "2022-12-19T17:03:35.323000",
              "content": "<p>me agree too</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2068137,
      "author_name": "Alex",
      "author_url": "",
      "post_date": "2022-12-17T14:55:31.883000",
      "content": "<p>Awesome tutorial 👍</p>\n<p>I think there's a little mistake in the ben_graham function.</p>\n<p>The image is not in grayscale here. You can therefore replace the color transform function or select only the value before blurring.</p>\n<pre><code> ():\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n  image = cv2.GaussianBlur(image, (,), )\n   image\n</code></pre>\n<p>or</p>\n<pre><code> ():\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n  image = cv2.GaussianBlur(image[], (,), )\n   image\n</code></pre>",
      "votes": 2,
      "replies": [
        {
          "id": 2068525,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-12-18T04:43:46.937000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2327695,
      "author_name": "kildongGo",
      "author_url": "",
      "post_date": "2023-07-03T05:47:15.520000",
      "content": "<p>Good Job ~!</p>\n<p>Realtime Histogram of Lab Colorspace is here: <a href=\"https://youtu.be/DXHI5xv6FF8\" target=\"_blank\">https://youtu.be/DXHI5xv6FF8</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2148812,
      "author_name": "bmjdoherty",
      "author_url": "",
      "post_date": "2023-02-17T17:28:38.823000",
      "content": "<p>This is incredibly useful. Thank you so much!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2134695,
      "author_name": "ZinkerovO",
      "author_url": "",
      "post_date": "2023-02-08T07:44:46.503000",
      "content": "<p>Thanks for your sharing. <br>\nI wonder whether the 'Augmentor' library could be used in this competition.<br>\nHave you tried that?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2085950,
      "author_name": "Abdoulaye Balde",
      "author_url": "",
      "post_date": "2023-01-04T13:34:05.903000",
      "content": "<p>A very nice tutorial thanks a lot for sharing this amazing material with us.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2070018,
      "author_name": "Sananda Chowdhury",
      "author_url": "",
      "post_date": "2022-12-19T14:35:11.963000",
      "content": "<p>The presentation is great</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2069804,
      "author_name": "Shivani P",
      "author_url": "",
      "post_date": "2022-12-19T10:40:57.070000",
      "content": "<p>Bookmarking this, really nice work <a href=\"https://www.kaggle.com/thedevastator\" target=\"_blank\">@thedevastator</a>!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2067634,
      "author_name": "@kaggleqrdl",
      "author_url": "",
      "post_date": "2022-12-16T23:27:06.043000",
      "content": "<p>Thanks, the post was very devastating!  :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2067401,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-12-16T17:06:38.960000",
      "content": "<p>Outstanding! If I could add more votes it would be at least 50 :) <br>\nBTW: nice avatar :) minimalistic. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2067421,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2022-12-16T17:21:07.233000",
      "content": "<p>Very nice thanks for sharing. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2133510,
      "author_name": "junseonglee11",
      "author_url": "",
      "post_date": "2023-02-07T13:07:51.523000",
      "content": "<p>Very useful content! Thank you so much.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2071353,
      "author_name": "Yahya Unlu",
      "author_url": "",
      "post_date": "2022-12-20T23:32:57.667000",
      "content": "<p>It's useful, Thank you for your effort.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2071146,
      "author_name": "Ashutosh Thakur",
      "author_url": "",
      "post_date": "2022-12-20T17:29:13.733000",
      "content": "<p>Thanks. This was helpful.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2070770,
      "author_name": "Akriti Upadhyay",
      "author_url": "",
      "post_date": "2022-12-20T10:01:52.680000",
      "content": "<p>This is helpful, thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2070587,
      "author_name": "priyanshu banerjee",
      "author_url": "",
      "post_date": "2022-12-20T07:18:26.093000",
      "content": "<p>Thanks fot the tips</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2068033,
      "author_name": "Khang Duong",
      "author_url": "",
      "post_date": "2022-12-17T12:47:39.033000",
      "content": "<p>Very useful !!! Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2067629,
      "author_name": "LeeLee24242424",
      "author_url": "",
      "post_date": "2022-12-16T23:11:56.293000",
      "content": "<p>Awesome, thanks.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2067389": "The following are image classification tips and tricks you might find useful for this competition. \n\nEnjoy! \n_____\n\n## Image Classification Tips & Tricks\n\n#### Augmentations\n\n#### Color Augmentations\n\n- **Color Skew:**\nThis augmentation randomly adjusts the hue, saturation, and brightness of the image by multiplying each channel by a randomly chosen coefficient. The coefficients are chosen from a range of [0:6;1:4] to ensure that the resulting image is not too distorted.\n\n```python\ndef color_skew(image):\n    h, s, v = cv2.split(image)\n    h = h * np.random.uniform(low=0, high=6)\n    s = s * np.random.uniform(low=1, high=4)\n    v = v * np.random.uniform(low=0, high=6)\n    return cv2.merge((h, s, v))\n```\n\n- **RGB Norm:**\nThis augmentation normalizes the RGB channels of the image by subtracting the mean value of each channel from the values in that channel and dividing by the standard deviation of the channel. This helps to standardize the values in the image and can improve the performance of the model.\n\n```python\ndef rgb_norm(image):\n    r, g, b = cv2.split(image)\n    r = (r - np.mean(r)) / np.std(r)\n    g = (g - np.mean(g)) / np.std(g)\n    b = (b - np.mean(b)) / np.std(b)\n    return cv2.merge((r, g, b))\n```\n\n- **Black and White:**\nThis augmentation converts the image to black and white by converting it to the grayscale color space.\n\n```python\ndef black_and_white(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n```\n\n- **Ben Graham: Greyscale + Gaussian Blur:**\nThis augmentation converts the image to grayscale and applies a Gaussian blur to smooth out any noise or details in the image.\n\n```python\ndef ben_graham(image):\n    image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n    image = cv2.GaussianBlur(image, (5, 5), 0)\n    return image\n```\n\n- **Hue, Saturation, Brightness:**\nThis augmentation converts the image to the HLS color space, which separates the image into its hue, saturation, and brightness channels.\n\n```python\ndef hsb(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2HLS)\n```\n\n- **LUV Color Space:**\nThis augmentation converts the image to the LUV color space, which is designed to be perceptually uniform and enables more accurate color comparison.\n\n```python\ndef luv(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2LUV)\n```\n\n- **Alpha Channel:**\nThis augmentation adds an alpha channel to the image, which can be used for transparency effects.\n\n```python\ndef alpha_channel(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2RGBA)\n```\n\t\n- **YZ Color Space:**\nThis augmentation converts the image to the XYZ color space, which is a device-independent color space that allows for more accurate color representation.\n\n```python\ndef xyz(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2XYZ)\n```\n\n- **Luma Chroma:**\nThis augmentation converts the image to the YCrCb color space, which separates the image into its luma (brightness) and chroma (color) channels.\n\n```python\ndef luma_chroma(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2YCrCb)\n```\n\n- **CIE Lab:**\n- This augmentation converts the image to the CIE Lab color space, which is designed to be perceptually uniform and enables more accurate color comparison.\n\n```python\ndef cie_lab(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2Lab)\n```\n\n- **YUV Color Space:**\nThis augmentation converts the image to the YUV color space, which separates the image into its luminance (brightness) and chrominance (color) channels.\n\n```python\ndef yuv(image):\n    return cv2.cvtColor(image, cv2.COLOR_RGB2YUV)\n```\n\n- **Center Crop:**\nThis augmentation randomly crops a rectangular area with an aspect ratio of [3/4,4/3], then randomly scales the crop by a factor between [8%,100%], and finally resizes the crop to an img_sizeXimg_size square. This is done randomly on each batch.\n\n```python\ntransforms.CenterCrop((100, 100))\n```\n\n**Flippings:**\nThis augmentation adds a probability of random horizontal flip to the image. For example, with a probability of 0.5, the image has a 50% chance of being horizontally flipped.\n\n```python\ndef flippings(image):\n    if np.random.uniform() < 0.5:\n        image = cv2.flip(image, 1)\n    return image\n```\n\n- **Random Crop:**\n\tThis augmentation randomly crops a rectangular area from the image.\n\n```python\ntransforms.RandomCrop((100, 100))\n```\n\n- **Random Resized Crop:**\nThis augmentation randomly resizes and crops a rectangular area from the image.\n\n```python\ntransforms.RandomResizedCrop((100, 100))\n```\n\n- **Color Jitter:**\nThis augmentation randomly adjusts the brightness, contrast, saturation, and hue of the image.\n\n```python\ntransforms.ColorJitter(brightness=0.5, contrast=0.5, saturation=0.5, hue=0.5)\n```\n\n- **Random Affine:**\nThis augmentation randomly applies an affine transformation to the image, which includes rotation, scaling, and shearing.\n\n```python\ntransforms.RandomAffine(degrees=45, translate=(0.1, 0.1), scale=(0.5, 2.0), shear=45)\n```\n\n- **Random Horizontal Flip:**\nrandomly flips the image horizontally with a probability of 0.5.\n\n```python\ntransforms.RandomHorizontalFlip()\n```\n\n- **Random Vertical Flip:**\nThis augmentation randomly flips the image vertically with a probability of 0.5.\n\n```python\ntransforms.RandomVerticalFlip()\n```\n\n- **Random Perspective:**\nThis augmentation randomly applies a perspective transformation to the image.\n\n```python\ntransforms.RandomPerspective()\n```\n\n- **Random Rotation:**\nThis augmentation randomly rotates the image by a given degree range.\n\n```python\ntransforms.RandomRotation(degrees=45)\n```\n\n- **Random Invert:**\nThis augmentation randomly inverts the colors of the image.\n\n```python\ntransforms.RandomInvert()\n```\n\n- **Random Posterize:**\nThis augmentation randomly reduces the number of bits used to represent each pixel value, resulting in a posterized effect.\n\n```python\ntransforms.RandomPosterize(bits=4)\n```\n\n- **Random Solarize:**\nThis augmentation randomly applies a solarize effect to the image, where pixels above a certain intensity threshold are inverted.\n\n```python\ntransforms.RandomSolarize(threshold=128)\n```\n\n- **Random Autocontrast:**\nThis augmentation randomly adjusts the contrast of the image by stretching the intensity values to the full available range.\n\n```python\ntransforms.RandomAutocontrast()\n```\n\n- **Random Equalize:**\nThis augmentation randomly equalizes the histogram of the image, resulting in increased contrast.\n\n```python\ntransforms.RandomEqualize()\n```\n\n#### Advanced Augmentations\n\n- **Auto Augment:**\nAuto Augment is an augmentation method that uses reinforcement learning to search for the optimal augmentation policies for a given dataset. It has been shown to improve the performance of image classification models.\n\n```python\nfrom autoaugment import AutoAugment\n\nauto_augment = AutoAugment()\nimage = auto_augment(image)\n```\n\n- **Fast Autoaugment:**\nFast Autoaugment is a faster implementation of the Auto Augment method. It uses a neural network to predict the optimal augmentation policies for a given dataset.\n\n```python\nfrom fast_autoaugment import FastAutoAugment\n\nfast_auto_augment = FastAutoAugment()\nimage = fast_auto_augment(image)\n```\n\n- **Augmix:**\nAugmix is an augmentation method that combines multiple augmented images to create a single, more diverse and realistic image. It has been shown to improve the robustness and generalization of image classification models.\n\n```python\nfrom augmix import AugMix\n\naug_mix = AugMix()\nimage = aug_mix(image)\n```\n\n- **Mixup/Cutout:**\nMixup is an augmentation method that combines two images by linearly interpolating their pixel values. Cutout is an augmentation method that randomly removes a rectangular region from the image. These methods have been shown to improve the robustness and generalization of image classification models.\n\n> - \"You take a picture of a cat and add some \"transparent dog\" on top of it. The amount of transparency is a hyperparam.\"\n\n```python\nx=lambda*x1+(1-lambda)x2\ny=lambda*x1+(1-lambda)y2\n```\n\n### Validation Time Augmentations\n\n* Adding validation set data augmentation will improve your model performance on test time augmentation and thus increase the absolution performance but if you seek to gain better performance use some of your validation set for data augmentation experiments or data augmentation experiments during training and use another part to increase your performance parameters, this will reduce your model overfitting.\n* You can use the test time augmentation by running the same model multiple times.\n* Run many validation data augmentation experiments at the same time to find the ones with the greatest impact your model performance.\n* **USE EXPERIMENTS:** Experiment. Experiment. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. \n\n### Test Time Augmentations\nAugmentations can be useful not juse during training but also during test time. \nSimply apply them on prediction and avg the results. \n\n- **Predict on different img scales:** Try to predict the same image but on different scales. \n- **Color Standardization:** Standardizatize as above.\n\n\n### Current High Scoring Models\n\nHere is a summary of the current highest scoring models on this competition:\n\n**tf_efficientnetv2_s**\nUsed on [this](https://www.kaggle.com/code/theoviel/rsna-breast-baseline-faster-inference-with-dali) notebook by [Theo Viel](https://www.kaggle.com/theoviel)\n\n**seresnext50_32x4d**\n- Find it on [this](https://www.kaggle.com/code/vslaykovsky/infer-pytorch-aux-targets-weighted-loss-thres) notebook by [Vladimir Slaykovskiy\n](https://www.kaggle.com/vslaykovsky)\n- Also used on [this](https://www.kaggle.com/code/hengck23/notebooke04a738685) notebook by [Heng](https://www.kaggle.com/hengck23)\n\n**tf_effv2_s_208_402**\n- Find it on [this](https://www.kaggle.com/code/radek1/fast-ai-starter-pack-train-inference) notebook by [Radek Osmulski](https://www.kaggle.com/radek1)\n\n\n### Models To Consider\n\nAnd some ideas and models to try out.\n\n**Swin Transformer**\n\n- Was the SOTA before EfficientNetV2, You can find example of it [here](https://www.kaggle.com/phalanx/train-swin-t-pytorch-lightning/notebook)\n\n**BeIT Transformer**\n\n- Also known to be a very strong model. \n\n**ViT Transformers**\n\n- This is the first vision transformer released. Still very powerful to this day.\n\n**EfficientNet L2**\n\n- The largest efficientnet V1, If you can somehow get it to work in terms of compute power it might yield high performance.\n\n### Hidden Layers After Your Backbone\n\nAdding more layers can be beneficial as you can use them to learn more advanced features, but it can also temper the fine-tuning of your large pretrained model.\n\n### Unfreeze Layer By Layer\n- A simple trick that can get you a tiny improvement would be to unfreeze layers of your pretrained backbone as the training progress. \n- **Adding more layers and freezing everything else:** As it turns out: Many solutions could be even further improved by incorporating another training phase after the pretrained model had been trained! This can be done by freezing the pretrained model and adding dense layers after it. You will get a tiny uplift in the upmost leaderboard with that change though. \n\n**Weight freezing and unfreezing in PyTorch**\n\n```python\n# Weight freezing\nfor param in model.parameters():\n  param.requires_grad = False # unfreeze weights in at all\n```\n\n```python\n# Weight unfreezing\nfor param in model.parameters():\n  param.requires_grad = True # unfreeze weights in at all\n```\n\n**Weight freezing and unfreezing in TensorFlow**\n\n```python\n# Weight freezing\nlayer.trainable = False\n```\n```python\n# Weight unfreezing\nlayer.trainable = True\n```\n\n### Learning-rates & LR Shedulers\n\nLearning rates and learning rate schedulers affect your models' training performance. Changing the learning rate can have a great impact on performance and also on training convergence. \n\n#### Learning rate schedulers\n\nIn recent times, One Cycle Cosine schedule has shown to provide better results on multiple NLP tasks, you can use it this way:\n\n**One Cycle Cosine scheduling in PyTorch**\n\n```python\nfrom torch.optim.lr_scheduler import CosineAnnealingLR\noptimizer = torch.optim.AdamW(optimizer_grouped_parameters, lr=args.learning_rate, eps=args.adam_epsilon)\nscheduler = CosineAnnealingLR(optimizer, T_max=num_train_optimization_steps)\nnum_training_steps = num_train_optimization_steps / args.gradient_accumulation_steps\n```\n\n```python\n# Update the scheduler\nscheduler.step()\n# step the learning rate scheduler here, you will want to step the learning rate scheduler only once per optimizer step nothing more nothing less. So in this case, it should be called before you expect the gradients to be applied.\n```\n\n**One Cycle Cosine scheduling in TensorFlow**\n\n```python\noptimizer = tf.keras.optimizers.Adam(learning_rate)\nscheduler = tf.keras.optimizers.schedules.CosineDecay(learning_rate, decay_steps=num_training_steps)\n```\n\n#### Some Learning Rate Schedulers Tips\n\n* Using “Triangular” or “One Cyclic” methods for learning rate scheduling can provide subtle but significant improvements - those intelligent methods of learning rate scheduling are can overcome some of the batch size issues and this will be obvious when using them, especially when plugging them into the multiple learning rate system. \n* Take the time to research for the best learning rate scheduling method for your task and the models you are using, it is a very important part of how your model will converge. \n* Learning rate schedules can be used to train models with lower batch sizes or multiple learning rates or any combination of those. \n* Try low learning rates first to see if a dramatic learning rate increase will help or hurt your performance. \n* Increasing the learning rate at later stages of training or multiple learning rates or high batch size or gradient accumulation or learning rate schedulers sometimes will help your model converge better, it is an advanced technique as sometimes it can hurt the performance but only if you give it too large a value - remember to test it. \n* Loss scaling can help reduce loss variance and improve gradient flow when using gradient accumulation or multiple learning rates or high batch size, but if you are trying to solve that problem by increasing batch size, try to increase the learning rate instead as it can sometimes yield better performance. \n* In recent times, one cycle learning rate scheduling approaches have taken the leading positions on public leaderboards and past winning solutions. \n\n#### Hyperparameters Of Optimizers\n\n> Crossposting from \"Transformers Hyperparams Tips\"\n\nThere are several things you need to know if you seek to get the best performance out of Adam optimizer:\n\n* Finding the best weight decay value can be tricky and experiments (and luck) will be your best friends here. \n* Another important hyperparameter is the `beta1` and `beta2` used in the Adam optimizer, choosing the best values depends on your task and the data. many new tasks can benefit from lower beta1 and higher beta2 while on established tasks they would perform the opposite. So again: Experiments will be your best friend here.\n* In the world of Adam optimizer, the number one rule is not to underestimate the importance of the optimizer epsilon value. The same principle of finding the best weight decay hyperparameter applies here.\n* Do not overuse the gradient clipping norm - it might sometimes be helpful when your gradients explode while the opposite is also true - it can prevent convergence on some tasks.\n* Gradient accumulation still sows to provide some subtle benefits, I usually accumulate gradients for about 2 steps but if your GPU does not run out of memory, you can push up to 8 gradient accumulation steps. Gradient accumulation is also useful when using mixed-precision. \n\n### Overfitting and regularization\n\n* Use dropout! adding layers dropouts between layers usually yields more training stability and more robust results, use dropouts with your hidden layers.\n* Dropout can also be used to increase performance by small margins, experiment with setting layer dropouts before training. the task and the model.\n* **If you are into regularization: **Regularization can provide marge uplift on performance when your NN is overfitting or underfitting, for normal Machine learning models, L1 or L2 regularization are OK but also try to use additive and hidden layers dropout.\n* **ALWAYS TEST OUT IDEAS USING EXPERIMENTS:** Use experiments. Experiment. Experiments and also try vanilla models as they still show subtle signs of performance. \n* **Multi Validations:** You can increase your model's robustness to overfitting by using multiple validations. This however comes at the cost of compute time.\n\n### Optimizer\nEveryone today are using Adam or AdamW. BUT something to consider is often times if you hyperparam-search SGD with momentum enough, you probably could get better results with it. But again this requires heavy tuning.\n\n**There are several notable optimizers that are important to know about:**\n- **AdamW:**  This is an extension to the Adam algorithm that prevents exponential weights decay of the model's weights in the outer layers as well as encourages penalized hyper volume below the default weights. \n- **Adafactor:** It was designed to have low memory usage and is scalable. This optimizer can offer significant optimizer performances using multiple GPUs (see below). \n- **Novograd:** Basically another Adam-like optimizer but with very better properties. It is one of the optimizers used by  to train the bert-large model. \n- **Ranger:** Ranger optimizer is a very interesting optimizer that has pretty good results in winning solutions it terms of performance optimization but it is obviously not very well known or supported so I leave that up to you to research. \n- **Lamb:** GPU optimized reusable Adam optimizer developed by the GLUE and QQP competition winner. \n- **Lookahead:** A popular optimizer you can use on top of other optimizers and it will provide you with some performance gains. \n\n### Label Smoothing\n\n> Original Paper: https://arxiv.org/pdf/1906.02629.pdf\n\nBasiclly a simple trick: y_true = y_true * (1.0 - ε) + 0.5 * ε [example: ε = 0.001]\nUsually works well. You can find it on some of the high scoring public notebook of this competition.\n\n\n**Tensorflow:** \n```python\nloss = BinaryCrossentropy(label_smoothing = label_smoothing)\n```\n\n**Pytorch:**\n\n```python\nfrom torch.nn.modules.loss import _WeightedLoss\n\nclass SmoothBCEwLogits(_WeightedLoss):\n    def __init__(self, weight = None, reduction = 'mean', smoothing = 0.0, pos_weight = None):\n        super().__init__(weight=weight, reduction=reduction)\n        self.smoothing = smoothing\n        self.weight = weight\n        self.reduction = reduction\n        self.pos_weight = pos_weight\n\n    @staticmethod\n    def _smooth(targets, n_labels, smoothing = 0.0):\n        assert 0 <= smoothing < 1\n        with torch.no_grad(): targets = targets * (1.0 - smoothing) + 0.5 * smoothing\n        return targets\n\n    def forward(self, inputs, targets):\n        targets = SmoothBCEwLogits._smooth(targets, inputs.size(-1), self.smoothing)\n        loss = F.binary_cross_entropy_with_logits(inputs, targets,self.weight, pos_weight = self.pos_weight)\n        if  self.reduction == 'sum': loss = loss.sum()\n        elif  self.reduction == 'mean': loss = loss.mean()\n        return loss\n```\t\t\n\n\n\n\n## Advanced Tricks\n\n### Knowledge Distillation\n> Use a large teacher network to guide the learning of a small network.\n\n- **Train the large model:** Train a large model on the data.\n- **Calculate soft target:** Use the trained large model to calculate soft target. That is, the output of the softmax after the large model is \"softened\"\n- **Studnet Model Training:** Train a student model based on the teachers output as an extra soft target loss function on the basis of the large model, and adjust the proportion of the two loss functions through interpulation.\n\n\n### Pseudo Labeling\n> Use a model to label unlabeled data (For example test data) then use the new labeled data you have for training your models.\n\n- **Train the teacher model:** Train a model on the data you have.\n- **Calculate soft target:** Use the trained large model to calculate soft target for unlabeled data.\n- **Important:** Use only the targets your model is \"sure\" about** Use only the highest confidence of the OOF predictions as pseudo labels to avoid mistakes as much as possible. (It might not work if you don't do this.).\n- **Studnet Model Training:** Train a student model on the new labeled data that you have.\n\n> **How do you CV pseudo labeling?** Simple: You concat the unlabeled data to your OOF and simply use them just as you would have used stacking. \n\n\n### Error Analysis\n\nAn important practice that can save you a lot of time is to use your model for finding harder or broken data samples.\nThere can be many reasons for an image to be \"harder\" for your model, for example, small target objects, different coloring, cut off target, invalid annotations and more.\n\n**Mistakes are Good News!**\n\nThese are exactly the samples that separate the top of the leaderboard from the rest of the participants.\nIf you have a hard time explaining what is going on with your model, it might be a good idea to take a look at the validation samples your model struggles with.\n\n### Finding Your Model's Errors\n\nThe easiest way to find errors is to sort validation samples by the model's confidence score and see which ones are being predicted with the lowest confidence.\n\n```python\nmistakes_idx = [img_idx for img_idx in range(len(train)) if int(pred[img_idx] > 0.5) != target[img_idx]]\nmistakes_preds = pred[mistakes_idx]\nsorted_idx = np.argsort(mistakes_preds)[:20]\n# Show the images of the sorted idx here..\n```\n\n\n",
    "2067445": "Of course, the above suggestion would be really helpful. But I've seen many similar post on each image competition. By doing so, it becomes easy for user to go from novice - grad-master tier in discussion. **Kaggle should place these common suggestion to their learning section** https://www.kaggle.com/learn/computer-vision.\n\n**If the tips and tricks are more oriented to a competition, only then it is purely useful.** For example, tips and tricks regarding of how to tackle imbalance dataset, or bigger input size training etc with limited resource - which are actual issues in this competition. And that also should be provided with some results and not some text from google. ",
    "2068137": "Awesome tutorial 👍\n\nI think there's a little mistake in the ben_graham function.\n\nThe image is not in grayscale here. You can therefore replace the color transform function or select only the value before blurring.\n\n```python\ndef ben_graham(image):\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2GRAY)\n  image = cv2.GaussianBlur(image, (5,5), 0)\n  return image\n```\nor\n\n```python\ndef ben_graham(image):\n  image = cv2.cvtColor(image, cv2.COLOR_RGB2HSV)\n  image = cv2.GaussianBlur(image[2], (5,5), 0)\n  return image\n```",
    "2327695": "Good Job ~!\n\nRealtime Histogram of Lab Colorspace is here: https://youtu.be/DXHI5xv6FF8",
    "2148812": "This is incredibly useful. Thank you so much!!",
    "2134695": "Thanks for your sharing. \nI wonder whether the 'Augmentor' library could be used in this competition.\nHave you tried that?",
    "2085950": "A very nice tutorial thanks a lot for sharing this amazing material with us.",
    "2070018": "The presentation is great",
    "2069804": "Bookmarking this, really nice work @thedevastator!",
    "2067634": "Thanks, the post was very devastating!  :)",
    "2067401": "Outstanding! If I could add more votes it would be at least 50 :) \nBTW: nice avatar :) minimalistic. ",
    "2067421": "Very nice thanks for sharing. ",
    "2133510": "Very useful content! Thank you so much.",
    "2071353": "It's useful, Thank you for your effort.",
    "2071146": "Thanks. This was helpful.",
    "2070770": "This is helpful, thanks for sharing",
    "2070587": "Thanks fot the tips",
    "2068033": "Very useful !!! Thanks for sharing",
    "2067629": "Awesome, thanks."
  }
}