Vision loss severely impacts object recognition and spatial cognition for limited vision individuals. It is a challenge to compensate for this using other sensory modalities, such as touch or hearing. This paper introduces StereoPilot, a wearable target location system to facilitate the spatial cognition of BVI. Through wearing a head-mounted RGB-D camera, the 3D spatial information of the environment is measured and processed into navigation cues. Leveraging spatial audio rendering (SAR) technology, it allows the navigation cues to be transmitted in a type of 3D sound from which the sound orientation can be distinguished by the sound localization instincts in humans. Three haptic and auditory display strategies were compared with SAR through experiments with three BVI and four sighted subjects. Compared with mainstream speech instructional feedback, the experimental results of the Fitts' law test showed that SAR increases the information transfer rate (ITR) by a factor of three for spatial navigation, while the positioning error is reduced by 40%. Furthermore, SAR has a lower learning effect than other sonification approaches such as vOICe. In desktop manipulation experiments, StereoPilot was able to obtain precise localization of desktop objects while reducing the completion time of target grasping tasks in half as compared to the voice instruction method. In summary, StereoPilot provides an innovative wearable target location solution that swiftly and intuitively transmits environmental information to BVI individuals in the real world.
To assist doctors in cardiovascular intervention, the use of robots in surgery can effectively separate interventionalists from X-ray radiation, prevent the spread of respiratory diseases, perform remote operations to alleviate the regional imbalance of medical resources and address some procedural challenges. However, a lack of force feedback in human computer interaction of current surgical robots affects surgical efficiency and safety. In cardiovascular surgery, the catheter bends and twists as it is the dynamic multipoint line contact with the vessel. It is difficult to accurately model and register it. In addition, time delay affects system transparency. This paper proposes a multi-information fuzzy fusion approach that is based on medical experience in order to predict position and orientation online. More specifically, Fitts law is developed with a two-motion collaborative to estimate the surgical movement time. Experimental results reveal that the proposed method can result in high system transparency in robot-assisted cardiovascular surgery. It could further improve accuracy and security in surgical procedures.
A hand-controller is a human-robot interactive device, which measures the 3-DOF (Degree of Freedom) position of the human hand and sends it as a command to control robot movement. The device also receives 3-DOF force feedback from the robot and applies it to the human hand. Thus, the precision of 3-DOF position measurements is a key performance factor for hand-controllers. However, when using a hybrid type 3-DOF hand controller, various errors occur and are considered originating from machining and assembly variations within the device. This paper presents a calibration method to improve the position tracking accuracy of hybrid type hand-controllers by determining the actual size of the hand-controller parts. By re-measuring and re-calibrating this kind of hand-controller, the actual size of the key parts that cause errors is determined. Modifying the formula parameters with the actual sizes, which are obtained in the calibrating process, improves the end position tracking accuracy of the device.
This paper presents a distributed robot sensor (DRS) system for force monitoring and control. The DRS system is an application of networked sensors that offers a reusable and portable design framework for Internet-based distributed measurement and control (DMC). It covers some important areas which play critical roles in future DMC development, including standardized transducer interface, open network communications, and Internet-based DMC application development. The portability and reusability is assured by the employment of IEEE 1451 smart transducer interface standards and object-oriented Java 2 platform. Experiment results have proved that the framework works well and is fully Web-enabled.
Multimodal fusion is essential for robots to fully perceive the external environment. Single modal information limits the ability of robots to recognize and grasp objects. Meanwhile, traditional cross-modal data generation methods produce poor images, resulting in the bad effects of multimodal fusion. To solve the problem of the poor image effects of multimodal generation and the lack of data for multimodal fusion, this study proposes a variational Bayesian Gaussian mixture-conditional generative adversarial network (BGM-CGAN) for generating diverse cross-modal noise data. With the variational Bayesian Gaussian mixture algorithm, a uniform distributed random noise group is generated into a single mixed variable. The generated mixed variable is then generated through a Gaussian mixture model to generate a series of Gaussian mixed noise groups. The generated mixed noise group is randomly selected into a single Gaussian noise and imported into a modal image and fused with the modal image to successfully generate a high-resolution heterogeneous modal image. The method restores the heterogeneous modal information and solves the problem of the insufficient information and poor quality of generated images under a single modal. In the end, we use a variety of evaluation indexes (IS, FID, SSIM, and PSNR) to compare the proposed BGM-CGAN and other algorithms for cross-modal image generation capabilities. Results show the effectiveness and feasibility of the proposed algorithm. In addition, the BGM-CGAN algorithm has wide application prospects and can be extended to cross-modal material retrieval, cross-modal texture recognition, and other fields.
Surface texture is one of the important cues for human beings to identify objects.Haptic texture measurement is necessary for object recognition by touch.This paper presents a novel design of a haptic texture sensor by imitating human active texture perception.A thin polyvinylidene fluoride (PVDF) film is used as the sensitive element to fabricate a high-accuracy, high-speed-response haptic texture sensor, and a mechanism is designed to produce the relative motion at a certain speed between the haptic texture sensor and the surface of the perceived object with constant contact force.Thus, the surface texture property can be measured as the output charge of the PVDF film of the sensor induced by the small height/depth variation of the moving object surface.The experiments reveal that the proposed active haptic sensor is effective in detecting the feature signals of surface texture, and the measurement signal can be used not only for the classification of the object surfaces, but also for haptic texture display in virtual reality.
Facial expression is an important carrier to reflect psychological emotion, and the lightweight expression recognition system with small-scale and high transportability is the basis of emotional interaction technology of intelligent robots. With the rapid development of deep learning, fine-grained expression classification based on the convolutional neural network has strong data-driven properties, and the quality of data has an important impact on the performance of the model. To solve the problem that the model has a strong dependence on the training dataset and weak generalization performance in real environments in a lightweight expression recognition system, an application method of confidence learning is proposed. The method modifies self-confidence and introduces two hyper-parameters to adjust the noise of the facial expression datasets. A lightweight model structure combining a deep separation convolution network and attention mechanism is adopted for noise detection and expression recognition. The effectiveness of dynamic noise detection is verified on datasets with different noise ratios. Optimization and model training is carried out on four public expression datasets, and the accuracy is improved by 4.41% on average in multiple test sample sets. A lightweight expression recognition system is developed, and the accuracy is significantly improved, which verifies the effectiveness of the application method.