You could ref:http://www.cvpapers.com/datasets.html
I'll paste that contents in the followings:
Participate in reproducible Detection PASCAL VOC, DataSet classification/detection competitions, segmentation competition, person Layout taster Competition Datasets LabelMe DataSet LabelMe is a web-based image annotation tool This allows researchers to label images and share T He annotations with the rest of the community. If you are using the database, we only ask this contribute to it, from time to time, by the with the labeling tool. BioId face Detection Database 1521 images with human faces, recorded under natural conditions, i.e. varying illumination a nd complex background. The eye positions has been set manually. CMU/VASC & PIE Face datasets Yale face DataSet Caltech Cars, Motorcycles, airplanes, Faces, Leaves, Backgrounds Caltech 101 Pictures of objects belonging to 101 categories Caltech $ Pictures of objects belonging to categories Daimler P Edestrian Detection Benchmark 15,560 Pedestrian and Non-pedestrian samples (image cut-outs) and 6744 additional full image s not containing pedestrians for bootstrappinG. The test set contains more than 21,790 images with 56,492 pedestrian labels (fully visible or partially occluded), Capt Ured from a vehicle in urban traffic. MIT pedestrian DataSet CVC pedestrian Datasets CVC pedestrian Datasets CBCL pedestrian Database mit face datasets Cbcl face Database mit car DataSet cbcl Car Database mit Street DataSet CBCL Street Database inria person Data Set A Large set of M arked up images of standing or walking people inria car dataset A set of car and non-car images taken in A parking lot NEA Rby inria inria Horse DataSet A set of horse and Non-horse images h3d dataset 3D skeletons and segmented regions for 1000 People in images HRI roadtraffic DataSet A Large-scale vehicle detection dataset Belgalogos 10000 images of natural scenes , with PNS different logos, and 2695 logos instances, annotated with a bounding box. Flickrbelgalogos 10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the Belga Logos DataSet. FlickrlogOs-32 the dataset FlickrLogos-32 contains photos depicting logos and is meant for the evaluation of Multi-Class logo Detec Tion/recognition as well as logo retrieval methods on Real-world images. It consists of 8240 images downloaded from Flickr. TME motorway Dataset 30000+ frames with vehicle rear annotation and classification (car and trucks) on Motorway/highway se Quences. Annotation semi-automatically generated using laser-scanner data. Distance estimation and consistent target ID over time available. Phos (color image database for illumination invariant feature selection) Phos is a Color image database of scenes Captu Red under different illumination conditions. More particularly, every scene of the database contains different Images:9 images captured under various strengths of Uniform illumination, and 6 images under different degrees of non-uniform illumination. The images contain objects of different shape, color and texture and can be used for illumination invariant featurE detection and selection. Californiand:an Annotated Dataset for near-duplicate Detection in Personal Photo collections california-nd contains 701 p Hotos taken directly from a real user ' s personal photo collection, including many challenging non-identical near-duplicate Cases, without the use of artificial image transformations. The dataset is annotated by ten different subjects, including the photographer, regarding near duplicates. USPTO algorithm challenge, detecting figures and part Labels in patents Contains drawing pages from US patents with manual LY labeled figure and part labels. Abnormal Objects Dataset Contains 6 Object categories similar to object categories in Pascal VOC that is suitable for Stu Dying the abnormalities stemming from objects. Human detection and tracking using rgb-d camera collected in a clothing store. Captured with Kinect (640*480, about 30fps) multi-task facial Landmark (MTFL) Datasets This dataset contains 12,995 face IM Ages collected from the Internet. THe images is annotated with (1) Five facial landmarks, (2) attributes of gender, smiling, wearing glasses, and head pose. Wider face:a face Detection Benchmark wider face dataset are A face Detection Benchmark datasets with images selected from The publicly available wider dataset. It contains 32,203 images and 393,703 face annotations. Classification PASCAL VOC, DataSet classification/detection competitions, segmentation competition, person Layout taster Competition Datasets Caltech Cars, Motorcycles, airplanes, Faces, Leaves, backgrounds Caltech 101 Pictures of objects belonging to 10 1 Categories Caltech Pictures of objects belonging to the categories Ethz Shape Classes A DataSet for testing object C Lass detection algorithms. It contains 255 test images and features five diverse shape-based classes (apple logos, bottles, giraffes, mugs, and Swans ). Flower Classification Data sets Flower Category DataSet Animals with attributes A datasets for Attribute Based classific ation. It consists of 30475 images of animals classes with six pre-extracted feature representations for each image. Stanford Dogs DataSet DataSet of 20,580 images of the dog breeds with bounding-box annotation, for fine-grained image Cate Gorization. Video Classification USAA DataSet The USAA DataSet includes 8 different semantic class videos which are HoMe videos of social occassions which feature activities of group of people. It contains around videos for training and testing respectively. Each video was labeled by attributes. The attributes can is broken down into five broad classes:actions, objects, scenes, sounds, and camera movement. McGill real-world Face Video Database This database contains 18000 video frames of 640x480 resolution from video Sequen CES, each of the which recorded from a different subject (female and male). Recognition Face and Gesture recognition Working group Fgnet face and Gesture recognition working Group fgnet Feret Face and Gesture Recognition working Group fgnet PUT face 9971 Images of all people labeled Faces in the Wild A database of face photograph s designed for studying the problem of unconstrained face recognition Urban scene recognition traffic Lights recognition, Lara ' s public benchmarks. Pubfig:public figures face Database The Pubfig database was a large, Real-world face dataset consisting of 58,797 images O F people collected from the Internet. Unlike existing face datasets, these images is taken in completely uncontrolled situations with Non-cooperativ E subjects. YouTube Faces The data set contains 3,425 videos of 1,595 different people. The shortest clip duration is frames, the longest clip was 6,070 frames, and the average length of a video clip is 181.3 Frames. Msrc-12:kinect gesture Data Set the Microsoft Cambridge-12 Kinect gesture data set consists OF Sequences of human movements, represented as Body-part locations, and the associated gesture to being recognized by the SYS Tem. Qmul underground re-identification (GRID) DataSet This dataset contains pedestrian image Pairs + 775 additional images Captured in a busy underground station for the "the" person re-identification. Person identification in TV series face tracks, features and shot boundaries from our latest CVPR. It is obtained from 6 episodes of Buffy the Vampire Slayer and 6 episodes of Big Bang theory. Chokepoint DataSet Chokepoint is a video dataset designed for experiments in person identification/verification under real -world surveillance conditions. The dataset consists of subjects (male and 6 female) in Portal 1 and subjects (both male and 6 female) in Portal 2. hieroglyph DataSet Ancient Egyptian hieroglyph DataSet. Rijksmuseum Challenge dataset:visual recognition for ART datasets over 110,000 photographic reproductions of the artworks ExhibiTed in the Rijksmuseum (Amsterdam, the Netherlands). Offers four automatic visual recognition challenges consisting of predicting the artist, type, material and creation year. Includes a set of baseline features, and offer an baseline based on State-of-the-art image features encoded with the Fishe R vector. The Ou-isir gait Database, treadmill Dataset treadmill gait datasets composed of subjects with 9 speed variations, s Ubjects with subjects, and 185 subjects with various degrees of gait fluctuations. The Ou-isir gait Database, Large Population Dataset Large Population gait datasets composed of 4,016 subjects. Pedestrian Attribute recognition at far Distance large-scale pedestrian Attribute (PETA) datasets, covering more than Tributes (e.g gender, age range, hair style, casual/formal) on 19000 images. Facescrub face DataSet The Facescrub dataset are a Real-world face datasets comprising 107,818 face images of 530 male and F Emale celebrities detected in images retrieved fromThe Internet. The images is taken under real-world situations (uncontrolled conditions). Name and gender annotations of the faces are included. Tracking Biwi Walking Pedestrians dataset Walking Pedestrians in busy scenarios from a bird eye view ' central ' pedestrian Crossing Sequences three pedestrian crossing sequences pedestrian Mobile Scene Analysis The set is recorded in Zurich, using a PA IR of cameras mounted on a mobile platform. It contains ' 298 annotated pedestrians in roughly 2 ' frames. Head tracking BMP image sequences. KIT AIS Dataset Data sets for tracking vehicles and people in aerial image sequences. MIT traffic Data set MIT traffic data set are for the activity analysis and crowded scenes. It includes a traffic video sequence of minutes long. It is recorded by a stationary camera. Shinpuhkan Dataset:multi-camera Pedestrian dataset for Tracking people across multiple cameras This dataset consists Of more than 22,000 images of people which is captured by cameras installed in a shopping mall "Shinpuh-kan". All images is manually cropped and resized to 48x128 pixels, GroupeD into Tracklets and added annotation. ATC Shopping Center Dataaset The tracking environment consists of multiple 3D range sensors, covering an area of about 900 M2, in the ' ATC ' shopping center in Osaka, Japan. Human detection and tracking using rgb-d camera collected in a clothing store. Captured with Kinect (640*480, about 30fps) multiple camera Tracking hallway corridor-multiple camera Tracking:an Indoo R Camera network DataSet with 6 cameras (contains ground plane homography). Multiple Object Tracking Benchmark A centralized Benchmark for Multi-Object Tracking. Segmentation Image segmentation with A bounding Box Prior DataSet Ground Truth database of images with:data, segmentation, Labelli Ng-lasso, Labelling-rectangle PASCAL VOC, DataSet classification/detection competitions, segmentation Competition, Person Layout Taster Competition Datasets Motion segmentation and Objcut Data cows for object segmentation, Five video SE Quences for motion Segmentation geometric Context Dataset Geometric context Dataset:pixel labels for seven geometric clas SES for images Crowd segmentation datasets This dataset contains videos of crowds and other high density moving objects . The videos is collected mainly from the BBC Motion Gallery and Getty Images website. The videos is shared only for the purposes. Consult the terms and conditions of use of these videos from the respective websites. Cmu-cornell icoseg Dataset Contains hand-labelled pixel annotations for groups of images, each group containing a Commo N Foreground. approximately IMAges per group, 643 images total. Segmentation evaluation database, gray level, images along with ground truth segmentations the Berkeley segmentation Dat Aset and Benchmark Image segmentation and Boundary detection. Grayscale and color segmentations for images, the images is divided into a training set of images, and a test set of images. Weizmann Horses 328 Side-view color images of horses that were manually segmented. The images were randomly collected from the WWW. Saliency-based video segmentation with sequentially updated priors videos as inputs, and segmented image sequences as G Round-truth Daimler Urban Segmentation DataSet The dataset consists of video sequences recorded in Urban traffic. The dataset consists of rectified stereo image pairs. Frames come with Pixel-level semantic class annotations into 5 classes:ground, building, vehicle, pedestrian, sky. Dense disparity maps is provided as a reference.Foreground/backgroundWallflower DataSet for evaluating background modelling algorithms foreground/background Microsoft Cambridge DataSet Foreg Round/background Segmentation and Stereo DataSet from Microsoft Cambridge Stuttgart Artificial Background subtraction Data Set the SABS (Stuttgart Artificial Background subtraction) DataSet is a Artificial dataset for pixel-wise evaluation of B Ackground models.saliency Detection (source)AIM IMAGES/20 observers (Neil D. B. Bruce and John K. Tsotsos 2005). Lemeu