Chapter 1 Introduction
1.1 Research Background and Significance
Energy drives the development of human society and the growth of the world economy, serving as the material foundation for survival across the globe. Coal, petroleum, natural gas, and other fossil fuels are currently the primary energy sources used by humanity. However, these fossil fuels are non-renewable resources with limited reserves on Earth, and their long-term consumption will eventually lead to energy depletion. Additionally, the extraction and utilization of these fossil fuels cause a series of environmental pollution issues. Coal mining loosens the surface soil and easily causes surface collapse; mine water discharge seriously pollutes groundwater resources and soil; direct combustion of fossil fuels produces large amounts of greenhouse gases such as carbon dioxide, triggering the greenhouse effect and causing severe consequences such as global temperature rise and warming. These issues not only endanger the balance of natural ecosystems but also threaten human survival. Therefore, facing the current situation of growing energy demand and severe pollution from fossil fuel energy, developing new clean energy sources and researching related utilization technologies hold important practical significance.
Solar energy, as a renewable energy foundation, can continuously provide light and heat, is safe to use, and does not pollute the environment, offering broad prospects for development and utilization. As a new energy source, solar power generation relies on two main approaches: solar thermal-electric conversion and photovoltaic (PV) power generation. The latter collects sunlight on solar cells and utilizes the photovoltaic effect to complete the conversion from light energy to electrical energy. Photoelectric conversion is safe, flexible, and is currently the primary method of using solar energy for power generation.
Traditional photovoltaic technology uses large-area silicon-based cells for power generation, resulting in low solar energy utilization. The extensive use of silicon raw materials makes it relatively expensive, with generation costs far exceeding traditional thermal power market prices. Concentrating photovoltaic (CPV) power generation technology is the third generation of PV power generation technology. It uses concentrating elements to focus large-area sunlight onto small-area solar cells, increasing the incident light intensity per unit area. This technology achieves higher photoelectric conversion efficiency. Concentrator elements are relatively low-cost, and the reduced solar cell area significantly lowers material usage, thus substantially decreasing the cost of CPV power generation.
CPV systems are classified by concentration ratio into low-concentration (below 100×), medium-concentration (100–300×), and high-concentration (above 300×) systems. A CPV system mainly consists of concentrators, solar cells, cooling components, and tracking control systems. The concentrator is the most critical component, determining the power generation performance of the entire system. CPV modules are manufactured using encapsulation integration technology, which encloses concentrators, solar cells, and heat-dissipation substrates in a module unaffected by external environmental conditions. Since the power generation efficiency of CPV systems largely depends on the total amount and intensity of sunlight concentrated onto the solar cells, the manufacturing of CPV modules demands high installation accuracy for both concentrators and solar cells.
Currently, few domestic manufacturers independently develop CPV module production lines; most rely on imported equipment. Therefore, researching precision photovoltaic panel visual alignment systems to achieve precise alignment and installation of concentrators and solar cells in CPV modules has important practical significance for the development of CPV technology.
1.2 Domestic and International Research Status
1.2.1 Research Status of Machine Vision
Machine vision technology refers to the technique of using machines to replace human eyes in providing visual functions in industry. With the increasing demand for product quality records and traceability documentation, machine vision has become an essential key technology in industrial automation and intelligence. A machine vision system captures real-object images through light sources, lenses, and industrial cameras, processes and analyzes these images using image processing algorithms in computers or smart cameras, and uses the extracted information to control and guide mechanical device operations. Machine vision involves knowledge from optics, image processing, motion control, and other fields, making it an interdisciplinary technology. Unlike computer vision, machine vision emphasizes solving practical industrial vision problems and requires substantial engineering application experience. Machine vision systems must adapt to harsh industrial production environments, possess high stability, precision, and fast processing speeds with good real-time performance.
Machine vision has entered industrial production lines due to several distinct characteristics: precision (high-resolution image acquisition devices far exceed human visual accuracy), non-contact operation (no direct contact with target objects), high speed (high-frame-rate cameras capture images on fast production lines, and high-performance processors complete image analysis in very short times), stability (avoids quality fluctuations caused by human physical and emotional conditions), and extended vision (recognizes broader spectral responses beyond visible light, including infrared, ultraviolet, and X-ray applications).
After more than sixty years of development, machine vision technology continues to evolve with new algorithms, technologies, and products. Many industries now recognize the potential of machine vision in improving product quality and production efficiency. Machine vision functions can be classified into target recognition, target positioning, target detection, and target measurement. Target recognition identifies objects based on special features such as characters, barcodes, and shapes. Target positioning obtains specific position information of measured objects to guide mechanical devices for subsequent processing or assembly. Target detection performs integrity checks on objects or determines whether defects exist. Target measurement measures geometric parameters of objects to ensure they remain within allowable tolerance ranges.
1.2.2 Research Status of Alignment Systems
Alignment systems are not limited to concentrating solar photovoltaic panel alignment; they also play important roles in PV cell printing, LCD alignment lamination, CD graphic printing, and electronic component mounting. In high-precision alignment applications, alignment systems use machine vision technology to obtain position information of aligned products and guide alignment mechanisms to achieve alignment.
The DEK Houyi series screen printing platform uses four CCD cameras and achieves alignment of chips or wafers with sizes between 125×125 mm and 165×165 mm, with maximum alignment accuracy reaching 12.5 μm and maximum production efficiency of 1,350 pieces per hour. The German EKRA SERIO 8000 series fully automatic screen printing system supports substrates up to 1000×610 mm. Its patented EVATM visual calibration system uses two high-resolution CCD cameras for image acquisition of the stencil and substrate, calculating alignment information through mark point position and spacing, achieving alignment accuracy up to 12.5 μm with production cycle times around 15 seconds.
Domestically, researchers have also studied alignment systems. He Zhenxing from South China University of Technology established error models for parallel platforms in solder paste printers, analyzed and researched systematic and random errors affecting image positioning results, and completed error compensation based on error analysis and test results. Liang Ruo from Huazhong University of Science and Technology designed an automatic alignment system for solar cell screen printing equipment. Zhang Zhiyao analyzed two main visual alignment methods used in LTCC industry printers. Miao Zhenhai proposed system calibration methods for FPC automatic loading machines and developed a visual inspection alignment system with correction accuracy reaching ±20 μm. Zhang Zhenya from Guangdong University of Technology developed an LCD alignment lamination system using LabVIEW software to adjust position deviations between tablet computer LCD displays and battery housings.
1.3 Project Source
This project originates from the cooperation project “Solar Module Visual Lamination System” between South China University of Technology and an industrial robotics company in Guangdong Province.
1.4 Main Research Content
To achieve precise alignment between lenses on the upper photovoltaic panel and solar cell components on the lower photovoltaic panel of concentrating solar photovoltaic panels, thereby increasing the photoelectric conversion efficiency of CPV modules, a visual alignment system for concentrating solar photovoltaic panels was developed. The system enables precise alignment between the upper and lower photovoltaic panels. The full text is divided into six chapters:
Chapter 1 introduces the research background and significance, followed by an overview of machine vision and alignment system research status worldwide.
Chapter 2 studies the alignment mechanism of the visual alignment system. The aligned product (CPV module) is introduced first, followed by the executive mechanism (alignment platform). The overall system scheme is then designed, and machine vision hardware selection is completed based on accuracy requirements and actual conditions.
Chapter 3 studies visual positioning detection algorithms for photovoltaic panels. The positioning detection algorithm for Mark points is determined, geometric features of Mark points are extracted through image processing, a shape-based template matching algorithm is created, and the algorithm is evaluated in terms of stability and real-time performance.
Chapter 4 completes four-camera image coordinate system calibration. Basic knowledge of planar coordinate system calibration is introduced, a calibration standard suitable for this system is designed, and the calibration algorithm between four camera image coordinate systems is detailed. Finally, software is written to complete calibration.
Chapter 5 completes alignment platform calibration. Kinematic modeling of the alignment platform is established first, followed by mathematical calibration using a three-point calibration method through software. When alignment results prove unsatisfactory, experimental calibration is ultimately adopted.
Chapter 6 integrates all algorithms and calibration data into the software and analyzes system operation results.
Chapter 2 Alignment Mechanism of the Visual Alignment System
2.1 Introduction to CPV Modules
The aligned product in this research is a refractive-type CPV module, as shown in the following figure. The CPV module is the most basic unit component of the CPV system, responsible for solar-to-electrical energy conversion. It consists of three parts: the upper photovoltaic panel, side frame, and lower photovoltaic panel.

The upper photovoltaic panel is a refractive-type lens plate that primarily focuses sunlight. It has multiple flat-plate point-focusing Fresnel lenses arranged on it. Fresnel lenses, also known as threaded lenses, have a smooth surface on one side and concentric circular rings of varying sizes on the other side, with each ring focusing incident light at the focal point. The lower photovoltaic panel is a solar cell panel with tempered glass as its main body, on which the same number of solar cell components as Fresnel lenses in the upper panel are distributed. A solar cell component consists of a bare cell and a cell substrate, with the bare cell welded at the center of the substrate. Components are connected in series through gold wires for current extraction. Additionally, an optical rod is typically installed on solar cell components to provide secondary concentration, improving light intensity distribution and increasing acceptance angle tolerance.
The side frame is placed between the upper and lower photovoltaic panels, providing support and protection. On one hand, it provides appropriate focal length for Fresnel lens concentration; on the other hand, it creates a sealed space to reduce external environmental interference with concentration and photoelectric conversion processes.
In this project, the Fresnel lenses in the upper photovoltaic panel have a concentration ratio of 750×, classifying them as high-concentration concentrators. The upper panel dimensions are 830×630×4 mm, and the lower panel dimensions are 830×630×3.2 mm, each weighing 3–5 kg. The alignment objective is to ensure that the focal point of each Fresnel lens in the upper panel falls within the bare cell of the corresponding solar cell component in the lower panel. Since the upper panel is manufactured by compression molding, which can only guarantee dimensional accuracy of each Fresnel lens and the spacing between lenses (but not the dimensional accuracy between peripheral Fresnel lenses and panel borders), edge-to-edge alignment cannot ensure that each Fresnel lens focal point aligns with its corresponding bare cell. This would seriously reduce module photoelectric conversion efficiency, necessitating the introduction of a visual alignment system.
2.2 Executive Mechanism of the Visual Alignment System
2.2.1 Introduction to the CPV Module Production Line
To improve production efficiency and product quality, CPV modules are manufactured on automated production lines. The production line primarily consists of industrial robots, conveyor belts, dispensing machines, and an alignment lamination mechanism. Industrial robots and conveyor belts handle photovoltaic panel loading/unloading, the dispensing machine applies adhesive around the panel edges, and the alignment lamination mechanism completes panel alignment and lamination with the side frame.
The alignment lamination mechanism’s executive components mainly consist of a lamination table and an alignment platform. Suction cups installed above the lamination table adsorb the upper panel onto the lower surface of the lamination table through vacuum suction. The lower panel is directly placed on the alignment platform.
The production process is as follows: two industrial robots (1 and 2) take the upper and lower panels from wooden boxes and place them on their respective conveyor belts. The conveyor belts transport panels to the dispensing station, where the dispensing machine applies adhesive around panel edges and sends completion signals to robots 3 and 4. These robots fetch the panels from the dispensing machine and simultaneously load them into the alignment lamination mechanism — the lamination table adsorbs the upper panel while the alignment platform carries the lower panel. After robots withdraw, the visual alignment system calculates position deviations through image acquisition and processing. Calibration algorithms convert position deviations into motor pulse values for the alignment platform, which then translates and rotates the lower panel to achieve correction. After alignment, robot 5 places the side frame onto the lower panel. The lamination table then moves toward the alignment platform, laminating the upper panel, side frame, and lower panel together. After curing the adhesive, robot 5 removes the finished CPV module to the product unloading conveyor.
2.2.2 Analysis and Introduction of the Alignment Platform
The alignment platform is one of the main research objects of this thesis. It is the primary executive mechanism for completing the photovoltaic panel alignment process. The alignment platform consists of a fixed platform and three executive motor devices. Each motor uses a ball screw-driven configuration with a push rod installed on the motor, and opposite each motor is a spring-loaded pull-rod mechanism. The fixed platform supports the lower panel, and vacuum devices on the fixed platform lift the panel during alignment to reduce friction. Motors and spring mechanisms fix the panel in position to prevent it from being blown off or running off, and then drive the push rods to translate and rotate the panel.
The alignment platform must perform X-direction translation, Y-direction translation, and angular rotation of the lower panel, constituting a planar three-degree-of-freedom alignment mechanism. Planar three-degree-of-freedom alignment mechanisms can be classified into serial and parallel types. Serial mechanisms have independent motions for each degree of freedom, making motion control simple and error compensation easy, but their stacked structure is bulky, causing error accumulation and swaying during rapid start-stop. It is also difficult for ordinary motors to achieve small-resolution rotation, and high-resolution motors are too expensive. Parallel mechanisms adopt a coplanar design philosophy, simplifying structure and reducing platform thickness, eliminating some errors and improving alignment accuracy. Parallel alignment mechanisms can easily achieve high-resolution rotation and are widely used in precision alignment applications.
The alignment platform used in this project is essentially a parallel alignment mechanism. Setting the line connecting the U and V motors as the X-axis direction and the line through the W motor perpendicular to the X-axis as the Y-direction: when U and V motors are stationary and the W motor moves, the platform moves in the X direction; when the W motor is stationary and U and V motors have equal movement, the platform moves in the Y direction; when all three move with U and V having equal movement, the platform translates in the XY plane; regardless of whether the W motor moves, when U and V motors have unequal movement, the platform rotates in the XY plane.
2.3 Overall Scheme of the Visual Alignment System
2.3.1 Vision Scheme
To achieve precise alignment between upper and lower photovoltaic panels, a vision scheme capable of obtaining precise position information of both panels is required. Visual alignment typically uses alignment reference Mark points to complete the alignment process. This project also adopts this approach. Mark points serve as reference points for lens and solar cell component positions, requiring high manufacturing precision. A photovoltaic panel’s position and posture are determined by two Mark points on its diagonal. For the upper panel, Mark points are the smallest circular teeth of Fresnel lenses in the first row, eleventh column and ninth row, second column. For the lower panel, Mark points are printed circles in black frames at diagonal boundary positions.
Since a single photovoltaic panel measures 830×630 mm — a relatively large size — using a single industrial camera to simultaneously image two Mark points would result in insufficient Mark point resolution in such a large field of view. Achieving high-precision positioning would require an ultra-high-pixel camera at excessive cost, and considering limited installation space, such an approach is clearly impractical. Large-field-of-view image acquisition typically uses camera arrays. Based on the positioning accuracy requirements, both upper and lower panel image acquisition devices use multi-camera array configurations. For a single photovoltaic panel, two cameras are used to acquire images of its two Mark points, meaning a total of four industrial cameras (arranged as two groups) complete Mark point image acquisition for both panels.
2.3.2 Control Scheme
The visual alignment system is applied to the alignment lamination mechanism, primarily completing position correction and calibration of the two photovoltaic panels, ensuring the final alignment deviation is less than 0.15 mm. The system consists of motion control and machine vision components. After robots place both panels into the alignment lamination platform, positions have certain randomness. To achieve precise alignment, the visual positioning system first acquires images of both panels, obtains panel position information through vision algorithms, feeds this back to the industrial control computer, which calculates position deviations and computes motor pulse values for the alignment platform based on calibration algorithms. The industrial control computer then drives executive motors through a motion control card to perform position correction.
The alignment process: after robots place the upper panel onto the lamination table with Mark points within camera fields of view, the lower panel is placed onto the alignment platform with Mark points within camera fields of view. Image acquisition and positioning of Mark points are performed, position deviations between both panels are calculated using Mark point positions and camera calibration algorithms, motor pulse amounts for the alignment platform are calculated through calibration algorithms, the industrial control computer drives motors through the motion control card, and the alignment platform pushes the lower panel to complete alignment.
2.4 Machine Vision Hardware Selection
Machine vision hardware is responsible for image acquisition, typically including industrial cameras, lenses, and lighting devices. The lighting device highlights essential target features, lenses produce clear images on camera sensors, and industrial cameras convert images into analog or digital video signals. The industrial control computer receives signals through camera interfaces and stores them in memory, completing image acquisition.
Industrial Camera: The camera’s role is to generate images from light focused on the image plane. Compared with ordinary cameras, industrial cameras offer higher image stability, higher transmission speeds, and stronger anti-interference capabilities. The most important component is the digital sensor, primarily CCD and CMOS types. CCD sensors excel in sensitivity, resolution, and image quality; CMOS sensors offer advantages in cost, power efficiency, and integration. Camera selection requires clarifying parameters including resolution, maximum frame rate, image mode, and interface type. Resolution determines image detail fineness. Resolution calculation requires field of view (FOV), feature resolution, and the number of pixels needed to represent feature resolution:
$$R_c = \frac{N_f \cdot FOV}{R_f}$$
For this system, the single-camera field of view references the 8×6 mm black frame of the lower panel Mark point. The vision accuracy requirement is 5 μm. Using one pixel per feature resolution yields a minimum camera resolution of 1600×1200. Frame rate refers to the number of images a camera captures per second. Since Mark point image acquisition occurs after panels are in place (static acquisition), high frame rates are unnecessary. Black-and-white cameras are selected since only contour features are needed. Gigabit Ethernet (GigE) interface cameras are chosen due to widespread Ethernet infrastructure, eliminating the need for additional image acquisition cards.
Based on the above analysis, four Basler acA1600-20gm industrial cameras were ultimately selected: resolution of 1626×1263, frame rate of 20 fps, monochrome, with GigE interface.
Lens: Common types include 360° optical imaging lenses, infrared optical lenses, fixed-focal-length lenses, and telecentric lenses. Fixed-focal-length lenses are most widely used in industry, operating on the pinhole imaging principle with perspective projection of objects. For the multi-camera array in this system, fixed-focal-length lenses would introduce perspective distortion due to inconsistent object distances among cameras, making it difficult to achieve parallel projection imaging. Telecentric lenses eliminate perspective distortion effects, maintaining constant magnification regardless of object distance within a certain range, producing parallel projection imaging. Therefore, telecentric lenses were ultimately selected.
Telecentric lens selection requires determining magnification and working distance. The Basler acA1600-20gm camera has a CCD target size of 7.2×5.4 mm. Selecting 0.8× magnification yields an object-side size of 9×6.75 mm, slightly larger than the required minimum field of view of 8×6 mm. Working distance determination must consider actual installation space. Given the narrow space between the alignment platform’s fixed platform and base, cameras cannot be installed vertically. Right-angle prisms are therefore introduced to deflect the optical path by 90 degrees. The working distance is 20 mm, the fixed platform thickness is 40 mm, with 5 mm reserved for camera position fine-tuning. Ultimately, a WWH-0865CT telecentric lens was selected with a working distance of 65 mm and 0.8× magnification. All four lenses use this same model. The upper camera group installs vertically, while the lower group installs laterally with right-angle prisms for optical path deflection.
Lighting Device: Light-emitting diodes (LEDs) are the most used light source in machine vision due to long service life, flash capability, easy brightness control through DC power, low power consumption, and availability in multiple colors. Various LED light sources exist for different applications. Initially, four white point light sources were used for forward illumination testing, capturing Mark point images of both panels. The upper panel Mark point image showed moderate contrast but clear contours; the lower panel, having a smooth bottom surface causing strong specular reflection, exhibited poor contrast with unclear boundaries between Mark point targets and background.
Since lower panel Mark point imaging quality is crucial, back-lighting was considered for contour acquisition. However, the mechanism above the fixed platform needs to accommodate both panels and the side frame for lamination, leaving no suitable installation position. Light sources were therefore installed on the lamination table, matching the upper camera group. Given the considerable distance light must travel through two panels, red light with the strongest penetrability was selected. Two high-power red ring lights illuminated from top to bottom. The resulting Mark point images showed significantly improved contrast with clear contours and distinct features compared with white point light source images.
The final machine vision hardware installation positions are: upper cameras above the lamination table, lower cameras below the alignment platform, red ring lights above the upper panel, and the industrial control computer processing image data.
Chapter 3 Research on Visual Positioning Detection Algorithms for Photovoltaic Panels
3.1 Determination of Positioning Detection Algorithm
3.1.1 Introduction to Mark Point Positioning Algorithms
In photovoltaic panel visual alignment, panel position information is determined by locating Mark points on the panels. Common Mark point shapes include circle, rectangle, triangle, diamond, cross, T-shape, and right angle. The accuracy of Mark point center positioning directly affects the overall panel alignment accuracy. Since Mark points on photovoltaic panels in this project are circular, several algorithms exist for circular Mark point positioning: Hough transform circle detection, edge detection fitting, Radon transform circle detection, and template matching.
Hough circle detection extends Hough transform, converting positioning problems from image space to parameter space. It finds parameter forms satisfying the majority of boundary points. Edge detection fitting primarily uses Blob analysis for coarse positioning, searches and marks Mark point regions based on connected component features, then extracts edges and fits them for positioning. Radon transform circle detection detects lines in parameter space: any two parallel lines through a circle have a distance equal to the diameter; the line parallel and equidistant between them passes through the center. Multiple such center lines intersect at the circle center. Template matching searches the image using templates, calculating similarity between template and image. It is broadly classified into gray-correlation-based matching and shape-based matching. The former uses gray values as similarity criteria; the latter uses pixel points, direction vectors, edges, and edge vectors as criteria.
3.1.2 Mark Point Positioning Algorithm Workflow
Before alignment, panels undergo edge dispensing. The upper panel Mark points are far from the dispensed area, unaffected by adhesive. However, lower panel Mark points are close to the adhesive area, sometimes causing interference and occlusion. Additionally, although robots place panels on the alignment platform, positioning deviations occasionally cause Mark points to not fully enter the field of view, resulting in missing Mark points. Since feed conditions are unstable, acquired Mark point images are complex and variable, demanding high algorithm robustness and precision. Hough transform has certain fault tolerance for interference or occlusion but is severely affected by numerous interferences or defects. Edge detection fitting requires high image stability and cannot handle such complex Mark point conditions with a single unified edge fitting algorithm. Radon transform requires complete circles without defects, otherwise defective edges affect positioning accuracy. Gray-correlation-based matching has poor anti-interference capability, prone to misidentification with nonlinear lighting changes, slight deformations, or occlusion. Shape-based matching uses edge features as similarity criteria; since edges are insensitive to nonlinear illumination changes, this algorithm has strong anti-interference ability. For both upper and lower panel Mark point images, the shape-based template matching algorithm was uniformly selected as the Mark point positioning algorithm. All positioning detection algorithms were developed on the Halcon platform.
The positioning algorithm workflow is: Mark point geometric feature extraction (ROI creation, image binarization, target region labeling, boundary edge fitting) and shape template creation.
3.2 Mark Point Geometric Feature Extraction
Image processing is performed on circular Mark point images to extract geometric features (circular contours) for shape template creation. To reduce complexity, one good-quality image of an upper panel Mark point and one of a lower panel Mark point (complete, unobstructed) are selected for contour extraction.
3.2.1 ROI Creation
Since Mark point regions occupy a relatively small proportion of acquired images, a region of interest (ROI) containing only the Mark point is created. Subsequent image processing is confined to this ROI, significantly reducing data volume and improving speed and accuracy. ROIs are created for both upper and lower panel Mark point images.
3.2.2 Image Binarization
Target Mark points and image backgrounds show distinct gray-level distributions, allowing global threshold segmentation. The best automatic global threshold algorithm is Otsu’s method. For a digital image of size M×N, with $n_i$ representing the number of pixels with gray level i, the probability of each gray level is:
$$p_i = \frac{n_i}{MN}$$
Setting threshold k divides image pixels into classes $T_1$ (gray levels $[0,k]$) and $T_2$ (gray levels $[k+1,255]$). The probabilities of pixels belonging to $T_1$ and $T_2$ are:
$$P_1(k) = \sum_{i=0}^{k} p_i, \quad P_2(k) = \sum_{i=k+1}^{255} p_i$$
The average gray values of classes $T_1$ and $T_2$ are:
$$f_1(k) = \frac{1}{P_1(k)} \sum_{i=0}^{k} ip_i, \quad f_2(k) = \frac{1}{P_2(k)} \sum_{i=k+1}^{255} ip_i$$
The overall average gray value is:
$$f_g = \sum_{i=0}^{255} ip_i = P_1(k) f_1(k) + P_2(k) f_2(k)$$
The between-class variance is:
$$\sigma^2 = P_1(k)(f_1 – f_g)^2 + P_2(k)(f_2 – f_g)^2$$
Otsu’s algorithm selects the k value maximizing $\sigma^2$ as the threshold for image segmentation. Applying this algorithm successfully binarizes Mark point images.
3.2.3 Target Region Labeling
The binarized image contains interference regions, and Mark point regions may contain pores. Hole-filling algorithms eliminate internal pores first. Then, an 8-connected method performs depth-first search and labels all connected components, obtaining all independent sub-regions. Mark point regions have relatively large areas compared with interference regions, so area features are used for screening. Based on camera resolution, field of view, and actual Mark point diameter, the area thresholds $A_{min}$ and $A_{max}$ are estimated to filter Mark point regions.
3.2.4 Boundary Edge Fitting
Morphological algorithms extract Mark point region boundaries. An appropriate 8-connected structuring element B performs dilation on Mark point region X:
$$X \oplus B = \{p \in \mathbb{Z}^2 : p = x + b, x \in X, b \in B\}$$
Dilation is iterative and increases region size. Then erosion is performed:
$$X \ominus B = \{p \in \mathbb{Z}^2 : p = x + b, b \in B\}$$
Erosion is the dual operation of dilation, reducing object size. Subtracting the eroded region F from the dilated region E produces the 8-connected Mark point boundary. Since boundaries contain many points and fitting a circle only requires the radius, edge data are fitted to circles. The least-squares circle fitting algorithm minimizes the sum of squared distances from all boundary points to the fitted circle:
$$\varepsilon^2 = \sum_{i=1}^{n} \left( r_i – c_i \right)^2 = \sum_{i=1}^{n} \left( \sqrt{(r_i-\alpha)^2 + (c_i-\beta)^2} – \rho \right)^2$$
where $(\alpha,\beta)$ is the center, $\rho$ is the radius, and $(r_i,c_i)$ are boundary points. Since Mark point boundaries have bumps and depressions, these outlier points affect the fitting accuracy. The Tukey weight function is introduced to reduce outlier influence:
$$\omega(\delta) = \begin{cases} \left(1 – (\delta/\tau)^2\right)^2 & |\delta| \leq \tau \\ 0 & |\delta| > \tau \end{cases}$$
where $\tau$ is the clipping factor and $\delta$ is the distance from boundary points to the fitted circle. After initial least-squares fitting, the Tukey weight function calculates weights iteratively. The fitted circles agree well with the circular Mark points, successfully extracting Mark point geometric features.
3.3 Shape Template Matching Algorithm
After extracting Mark point geometric features, a shape-based template matching algorithm is constructed. Edge extraction methods compute edge point coordinates and gradient directions for Mark point edges and search images. The shape template is defined as point set $p_i = (r_i, c_i)^T$ with direction vectors $d_i = (t_i, u_i)^T$, and the image is defined as point set $q = (r, c)^T$ with direction vectors $e_{r,c} = (v_{r,c}, w_{r,c})^T$. Using the linear transformation model $p’_i = Ap_i$ and $d’_i = (A^{-1})^T d_i$, the sum of normalized direction vector dot products between template and corresponding image points is the similarity measure:
$$s = \frac{1}{n} \sum_{i=1}^{n} \frac{d_i^{‘T} e_{q+p’_i}}{\|d’_i\| \cdot \|e_{q+p’_i}\|} = \frac{1}{n} \sum_{i=1}^{n} \frac{t’_i v_{r+r’_i,c+c’_i} + u’_i w_{r+r’_i,c+c’_i}}{\sqrt{t_i^{‘2} + u_i^{‘2}} \cdot \sqrt{v_{r+r’_i,c+c’_i}^2 + w_{r+r’_i,c+c’_i}^2}}$$
Since direction vectors are normalized, similarity is unaffected by occlusion, clutter, and nonlinear illumination changes. For different applications, different similarity functions are used. When template and image have the same light-dark contrast direction, the above formula is used. When contrast direction is reversed:
$$s_2 = \frac{1}{n} \sum_{i=1}^{n} \frac{|d_i^{‘T} e_{q+p’_i}|}{\|d’_i\| \cdot \|e_{q+p’_i}\|}$$
For cases ignoring local contrast direction changes:
$$s_3 = \frac{1}{n} \sum_{i=1}^{n} \frac{|d_i^{‘T} e_{q+p’_i}|}{\|d’_i\| \cdot \|e_{q+p’_i}\|}$$
Since Mark point images are extremely unstable, Eq. (s₃) is ultimately selected to maximize search capability. All similarity values are less than or equal to 1; a value of 1 indicates identical template and target. During search, a threshold $s_{min}$ is set to compare with similarity, stopping calculations when partial sums fall below thresholds. Setting:
$$s_j = \frac{1}{n} \sum_{i=1}^{j} \frac{d_i^{‘T} e_{q+p’_i}}{\|d’_i\| \cdot \|e_{q+p’_i}\|}$$
The remaining $n-j$ edge points have dot product sums bounded by $(n-j)/n = 1 – j/n$. Therefore, when $s_j < s_{min} – 1 + j/n$, the similarity cannot reach the threshold, and calculation stops early. Due to occasional occlusion and interference, $s_{min}$ is set to 0.5.
The computational complexity $O(whn)$ depends on image width w, height h, and template edge point number n. Computing similarity across the entire image is time-consuming. To ensure real-time performance, a coarse-to-fine strategy uses image pyramids — repeatedly downsampling template and image by 2, stacking from largest to smallest, forming a pyramid structure. A 2×2 mean filter smooths images without frequency response issues or image translation. Pyramid search first computes similarity at the top level, identifies potential match positions with local maximum similarity values, then tracks these positions down pyramid levels to the first level. Each additional pyramid level speeds up matching by 16 times. Too many levels damage image information; the level count depends on target feature characteristics, ensuring features remain identifiable at the highest level. For both upper and lower Mark point images, level 6 significantly degrades circular Mark point identifiability, so the pyramid level is set to 5.
Using the shape template matching algorithm to search original Mark point images successfully found Mark point positions.
3.4 Evaluation of Mark Point Positioning Algorithm
3.4.1 Stability of the Positioning Algorithm
Stability primarily tests the algorithm’s ability to successfully search targets under linear, nonlinear, noise, interference, partial missing, or occlusion conditions. Matching results were evaluated for images with low contrast, partially missing Mark points, and interference/occlusion. The algorithm demonstrated good robustness and stability.
3.4.2 Runtime of the Positioning Algorithm
For convenient software integration, shape templates are saved in .shm format after creation. The matching runtime was measured on a dual-core 2.60 GHz CPU, 4 GB RAM, with target images of 1626×1236 pixels. Upper panel Mark point matching averaged approximately 5 ms; lower panel Mark points, being larger with more template edge points and higher computational complexity, averaged approximately 20 ms. Including image acquisition time, individual Mark point positioning takes less than 100 ms, satisfying video positioning system real-time requirements.
Table 3-1 Positioning Algorithm Runtime
| Experiment | Upper Panel Mark Point [ms] | Lower Panel Mark Point [ms] |
|---|---|---|
| 1 | 4.760 | 19.174 |
| 2 | 5.190 | 19.588 |
| 3 | 5.131 | 21.173 |
| 4 | 5.025 | 20.343 |
| 5 | 4.867 | 20.314 |
| 6 | 5.083 | 19.706 |
| 7 | 5.035 | 21.230 |
| 8 | 4.943 | 21.696 |
| 9 | 4.895 | 19.755 |
| 10 | 4.380 | 21.129 |
Chapter 4 Four-Camera Image Coordinate System Calibration
4.1 Planar Coordinate System Calibration Principles
After completing Mark point positioning detection, Mark point position information can be obtained from four camera images. However, the four cameras are independent, each having its own image coordinate system. The obtained Mark point positions are positions in four independent image coordinate systems, making it impossible to directly calculate position deviations between upper and lower panels. Therefore, calibrating the four camera image coordinate systems requires calculating transformation relationships between each coordinate system.
Image coordinate systems are two-dimensional planar coordinate systems. Coordinate system calibration calculates transformations from one coordinate system to another. A coordinate system is described by position and orientation: position uses a position vector describing the origin in the plane, and orientation uses a rotation matrix. For coordinate system $\{B\}$ relative to $\{A\}$, the position vector and rotation matrix are:
$$^A P_{BORG} = \begin{bmatrix} p_x \\ p_y \end{bmatrix}$$
$$^A_B R = \begin{bmatrix} r_{11} & r_{12} \\ r_{21} & r_{22} \end{bmatrix} = \begin{bmatrix} ^A \hat{X}_B & ^A \hat{Y}_B \end{bmatrix}$$
When two planar coordinate systems have the same orientation, only translation differs. Given point P in coordinate system $\{B\}$, its position in $\{A\}$ is:
$$^A P = ^B P + ^A P_{BORG}$$
When origins coincide, only rotation differs:
$$^A P = ^A_B R ^B P$$
When both orientation and origin translation differ:
$$^A P = ^A_B R ^B P + ^A P_{BORG}$$
Expanding:
$$\begin{bmatrix} X \\ Y \end{bmatrix} = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} \begin{bmatrix} X \\ Y \end{bmatrix} + \begin{bmatrix} q_x \\ q_y \end{bmatrix}$$
where $\theta$ is the rotation angle between coordinate systems and $q_x$, $q_y$ are translation vector components. Calibration between planar coordinate systems requires obtaining the rotation angle and translation quantities. Using the strategy of calibrating pixel-physical size ratios with standard parts, the calibration standard method is also used for image coordinate systems.
4.2 Design of the Calibration Standard
4.2.1 Calibration Benchmark Analysis
The calibration standard must satisfy two requirements: 1) have feature points with precise physical dimensions; 2) have multiple feature points allowing calculation of rotation angles and translations through their relationships.
Using the photovoltaic panel product itself was considered first. Mark point manufacturing precision is extremely high and physical distances are known, but each Mark point image contains only one Mark point, making it impossible to calculate angle information from a single point.
Calibration boards (circular dot and checkerboard types) were then considered. These use centroid information or corner points for calibration. Manufacturing precision is high, and large boards can calibrate camera coordinate system relationships. However, due to the significant height difference between upper and lower camera installation positions, combined with limited lens depth of field, upper and lower cameras cannot focus on the same plane. A calibration board can only calibrate among upper cameras or among lower cameras, not between upper and lower cameras. Due to these special spatial position characteristics, a dedicated calibration standard was designed and manufactured.
4.2.2 Calibration Standard Design
The calibration standard must have sufficient thickness so both upper and lower cameras can focus on its feature points. The thickness is designed by referencing the side frame thickness. Since cameras are installed at panel diagonal positions, the calibration standard need not be a large flat block like a panel; only the diagonal portions are needed, but flatness must be high. Feature point positions reference upper and lower panel Mark point positions to ensure features enter all camera fields of view. Multiple feature points must appear in each camera image to determine angle and distance information.
Distance information is obtained from feature point spacing. Although angle information can be determined from two points, manufacturing errors could cause significant angle errors. Increasing feature point quantity and using linear regression to find the line closest to all points approximates the theoretical design line. More points yield more accurate results. Considering manufacturing cost and limited field of view, three feature points are chosen for determining lines. To simplify subsequent calibration calculations, feature points in the same plane groups captured by camera pairs must also be collinear.
Feature point distribution on the calibration standard: circular holes represent feature points. Given each camera field of view is approximately 9×6.75 mm, feature point holes are designed with 1.5 mm diameter and 2.5 mm spacing between holes within groups. Material selection is mold steel for sufficient strength and hardness. Technical requirements include smooth bright surfaces, no scratches, no distortion or deformation of the main body, sharp edges deburred, and no chamfering on hole edges.
4.3 Four-Camera Image Coordinate System Calibration Algorithm
4.3.1 Unifying Image Coordinate System Directions
Upper and lower cameras have different installation methods, so image coordinate system directions differ. To simplify calibration calculations, all camera image coordinate system directions are unified first. Since alignment references the upper panel position, the upper-left camera image coordinate system is selected as the reference coordinate system. Other camera coordinate systems rotate to match: upper-right rotates 180°, lower-left rotates −90°, lower-right rotates 90°, with counterclockwise as positive.
4.3.2 Upper-Right and Upper-Left Camera Image Coordinate System Calibration
The upper-left and upper-right cameras simultaneously capture the calibration standard. Points 1, 2, 3 (top to bottom) in the upper-left image have coordinates $(X_1,Y_1)$, $(X_2,Y_2)$, $(X_3,Y_3)$; points 4, 5, 6 in the upper-right image have coordinates $(x_4,y_4)$, $(x_5,y_5)$, $(x_6,y_6)$. For each group, a univariate linear regression line is found to minimize the sum of vertical distances from scattered points to the line. For upper-left points, the regression line is $y = kx + b$. Mean values are:
$$\bar{X} = \frac{1}{n}\sum_{i=1}^{n} X_i, \quad \bar{Y} = \frac{1}{n}\sum_{i=1}^{n} Y_i$$
Coefficients are calculated:
$$L_{XX} = \sum_{i=1}^{n}(X_i-\bar{X})^2, \quad L_{YY} = \sum_{i=1}^{n}(Y_i-\bar{Y})^2, \quad L_{XY} = \sum_{i=1}^{n}(X_i-\bar{X})(Y_i-\bar{Y})$$
The regression line slope and intercept:
$$k = \frac{L_{XY}}{L_{XX}}, \quad b = \bar{y} – k\bar{x}$$
The angle between the regression line and the upper-left camera image coordinate system is:
$$\alpha_1 = \arctan(k)$$
Similar calculations for upper-right points yield angle $\alpha_2$. If camera orientations are identical, $\alpha_1$ equals $\alpha_2$. Therefore, the rotation angle between upper-right and upper-left camera coordinate systems:
$$\theta = \alpha_2 – \alpha_1$$
For translation vector calculation, the middle point of each group (points 2 and 5) is used. With pixel equivalent K and actual physical distance l between points 2 and 5, the pixel distance L = Kl. The coordinates of point 5 in the upper-left camera coordinate system are $(X_5, Y_5)$:
$$X_5 = X_2 + L\cos\alpha_1, \quad Y_5 = Y_2 – L\sin\alpha_1$$
Using the transformation:
$$\begin{bmatrix} X_5 \\ Y_5 \end{bmatrix} = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} \begin{bmatrix} x_5 \\ y_5 \end{bmatrix} + \begin{bmatrix} T_X \\ T_Y \end{bmatrix}$$
Solving simultaneously yields:
$$T_X = X_2 + L\cos\alpha_1 – x_5\cos\theta + y_5\sin\theta$$
$$T_Y = X_2 – L\sin\alpha_1 – x_5\sin\theta – y_5\cos\theta$$
This completes the upper-right and upper-left camera image coordinate system calibration.
4.3.3 Lower-Left and Upper-Left Camera Image Coordinate System Calibration
The upper-left and lower-left cameras simultaneously capture the calibration standard. Points 1, 2, 3 in the upper-left image have coordinates $(X_1,Y_1)$, $(X_2,Y_2)$, $(X_3,Y_3)$; points 7, 8, 9 in the lower-left image have coordinates $(x_7,y_7)$, $(x_8,y_8)$, $(x_9,y_9)$. Univariate linear regression lines are established, yielding angles $\alpha_1$ and $\alpha_3$ between regression lines and respective coordinate systems. The collinear lines of these two point groups intersect at point O in the XY plane with an angle $\gamma$ between them.
The rotation angle between coordinate systems:
$$\theta = \alpha_3 – (\alpha_1 + \gamma)$$
For translation vector calculation, middle points 2 and 8 are used. With actual distances $l_1$ between point 2 and O and $l_2$ between point 8 and O, pixel distances are $L_1 = Kl_1$ and $L_2 = Kl_2$. Point 8 coordinates in the upper-left camera coordinate system:
$$X_8 = X_2 + L_1\cos\alpha_1 – L_2\cos\beta$$
$$Y_8 = Y_2 + L_1\sin\alpha_1 – L_2\sin\beta$$
where $\beta = \frac{\pi}{2} – \alpha_1 – \gamma$. Using the transformation relationship:
$$\begin{bmatrix} X_8 \\ Y_8 \end{bmatrix} = \begin{bmatrix} \cos\theta & -\sin\theta \\ \sin\theta & \cos\theta \end{bmatrix} \begin{bmatrix} x_8 \\ y_8 \end{bmatrix} + \begin{bmatrix} T_X \\ T_Y \end{bmatrix}$$
Solving simultaneously yields:
$$T_X = X_2 + L_1\cos\alpha_1 – L_2\cos\beta – x_8\cos\theta + y_8\sin\theta$$
$$T_Y = X_2 + L_1\sin\alpha_1 – L_2\sin\beta – x_8\sin\theta – y_8\cos\theta$$
4.3.4 Lower-Right and Upper-Left Camera Image Coordinate System Calibration
Based on calibration standard feature distribution, lower-right and lower-left camera calibration uses the same algorithm as upper-right and upper-left camera calibration. Lower-right to upper-left calibration is achieved segment-wise: first calibrating lower-right to lower-left, then using the lower-left to upper-left calibration result.
4.4 Four-Camera Image Coordinate System Calibration Software
To ensure convenient calibration and avoid manual calculation truncation and rounding errors, the calibration algorithm was programmed using C#. The software includes capture, positioning, calculation, and save components. The calibration process is as follows: the calibration standard is placed on the alignment lamination mechanism, the lamination table moves to the lamination position, and the calibration standard position is fine-tuned so all feature points appear in all camera fields of view. Using the software’s capture button, four cameras perform triggered image acquisition. Images are rotated to unify coordinate system directions and displayed. The positioning button calls the positioning algorithm (shape-based template matching) to locate feature points, obtaining coordinates of three feature points in each camera image. Linear regression algorithms calculate angles. The calculate button invokes the respective calibration algorithms using feature point coordinates, regression line angles, and actual calibration standard feature positions. The save button stores all calibration data in .ini format for subsequent queries and calls.
Chapter 5 Alignment Platform Calibration
5.1 Kinematic Modeling of the Alignment Platform
After four-camera image coordinate system calibration, transformation relationships between all four camera image coordinate systems are established. Mark point positions of both panels can be unified into the upper-left camera coordinate system, where position deviations between panels are calculated. However, position correction ultimately occurs on the alignment platform, so the position deviation must be transformed from the upper-left camera image coordinate system to the alignment platform coordinate system.
This system focuses on static position and posture, so only kinematic analysis is performed. Given mechanism input parameters, solving mechanism position and posture is the forward kinematics problem; conversely, given position and posture, solving input parameters is the inverse kinematics problem. This system requires solving motor input pulses given the final position and posture — an inverse kinematics problem.
Using the geometric method for kinematic modeling: coordinate system xy represents the final position and posture of the alignment platform; coordinate system XY represents the initial position and posture. Knowing X-direction deviation $\Delta x$, Y-direction deviation $\Delta y$, and rotation deviation $\Delta\theta$, the input quantities $\Delta u$, $\Delta v$, $\Delta w$ for U, V, W motors are solved. $L_1$, $L_2$, $L_3$ represent installation positions of U, V, W motors — distances from each motor axis to the alignment platform coordinate system origin O. According to geometric relationships:
$$\Delta u = \Delta y + (L_1 – \Delta x)\tan\Delta\theta$$
$$\Delta v = \Delta y + (L_2 – \Delta x)\tan\Delta\theta$$
$$\Delta w = \Delta x – (L_3 – \Delta y)\tan\Delta\theta$$
During alignment, panel deviation does not equal platform translation. In coordinate systems, point deviation does not equal origin deviation. A relationship model between panel deviation and platform translation is established. Let $(x_0, y_0)$ be the midpoint of lower panel Mark points in the platform coordinate system and $(x_1, y_1)$ be the midpoint of upper panel Mark points. The angle deviation between lower and upper panels is $\Delta\theta$, which is the same as the platform coordinate system deviation. The alignment process makes point $(x_0, y_0)$ coincide with point $(x_1, y_1)$ while rotating the platform coordinate system by $\Delta\theta$.
According to planar coordinate system calibration:
$$\begin{bmatrix} \Delta x \\ \Delta y \end{bmatrix} = \begin{bmatrix} \cos\Delta\theta & \sin\Delta\theta \\ -\sin\Delta\theta & \cos\Delta\theta \end{bmatrix} \begin{bmatrix} x_1 – x_0 \\ y_1 – y_0 \end{bmatrix}$$
Solving:
$$\Delta x = (x_1 – x_0)\cos\Delta\theta + (y_1 – y_0)\sin\Delta\theta$$
$$\Delta y = -(x_1 – x_0)\sin\Delta\theta + (y_1 – y_0)\cos\Delta\theta$$
Substituting into previous equations yields the three motor displacements.
5.2 Mathematical Calibration of the Alignment Platform
5.2.1 Three-Point Calibration Method Principle
After kinematic modeling, motor displacements can be calculated from platform coordinate system deviations. However, obtaining panel deviations requires calibrating the upper-left camera image coordinate system with the platform coordinate system. The three-point calibration method is used: three non-collinear points calculate the transformation between two coordinate systems. Assuming camera image coordinate system is parallel to platform coordinate system, for a point at $(x,y)$ in the camera coordinate system and $(X,Y)$ in the platform coordinate system:
$$\begin{bmatrix} X \\ Y \end{bmatrix} = \begin{bmatrix} A & B \\ C & D \end{bmatrix} \begin{bmatrix} x \\ y \end{bmatrix} + \begin{bmatrix} E \\ F \end{bmatrix}$$
Since image coordinates are in pixels and platform coordinates are in motor pulses, units differ, so scale coefficient K is incorporated into the matrix coefficients. Let three non-collinear points have camera coordinates $(x_1,y_1)$, $(x_2,y_2)$, $(x_3,y_3)$ and platform coordinates $(X_1,Y_1)$, $(X_2,Y_2)$, $(X_3,Y_3)$. Substituting all three points:
$$X_i = Ax_i + By_i + E$$
$$Y_i = Cx_i + Dy_i + F$$
Subtracting point 1 equations from points 2 and 3:
$$X_2 – X_1 = A(x_2-x_1) + B(y_2-y_1)$$
$$X_3 – X_1 = A(x_3-x_1) + B(y_3-y_1)$$
$$Y_2 – Y_1 = C(x_2-x_1) + D(y_2-y_1)$$
$$Y_3 – Y_1 = C(x_3-x_1) + D(y_3-y_1)$$
Solving these equation systems:
$$A = \frac{(X_2-X_1)(y_3-y_1) – (X_3-X_1)(y_2-y_1)}{(x_2-x_1)(y_3-y_1) – (x_3-x_1)(y_2-y_1)}$$
$$B = \frac{(X_2-X_1)(x_3-x_1) – (X_3-X_1)(x_2-x_1)}{(x_2-x_1)(y_3-y_1) – (x_3-x_1)(y_2-y_1)}$$
$$C = \frac{(Y_2-Y_1)(y_3-y_1) – (Y_3-Y_1)(y_2-y_1)}{(x_2-x_1)(y_3-y_1) – (x_3-x_1)(y_2-y_1)}$$
$$D = \frac{(Y_2-Y_1)(x_3-x_1) – (Y_3-Y_1)(x_2-x_1)}{(x_2-x_1)(y_3-y_1) – (x_3-x_1)(y_2-y_1)}$$
Then:
$$E = X_1 – Ax_1 – By_1$$
$$F = Y_1 – Cx_1 – Dy_1$$
5.2.2 Alignment Platform Calibration Software
Using the three-point calibration method, three non-collinear points with both camera image coordinates and platform coordinates are needed. The alignment approach is: using the lower panel, the platform moves three times, ensuring U and V motors have equal displacement and the three positions are non-collinear. The motor pulse values serve as platform coordinates $(X_1,Y_1)$, $(X_2,Y_2)$, $(X_3,Y_3)$; the lower-left camera captures Mark points three times, obtaining camera coordinates $(x’_1,y’_1)$, $(x’_2,y’_2)$, $(x’_3,y’_3)$. Using the lower-left-to-upper-left calibration relationship, these coordinates are converted to the upper-left camera coordinate system, then the three-point calibration algorithm calibrates the platform.
The software was programmed in C#. The calibration process: place the lower panel on the alignment platform, ensuring Mark points are within the lower-left camera field of view. For the first point, click capture to acquire the first Mark point image, use shape-based template matching to locate it, record motor pulse values. For the second point, click any direction button in the motor fine-tune group, platform moves 400 pulses in that direction, capture, locate, and record. For the third point, click a direction ensuring non-collinearity with previous two points, move 400 pulses, capture, locate, and record. Clicking calculate invokes the calibration relationship to convert the three points from lower-left to upper-left coordinate systems, then the three-point calibration algorithm is applied. The save button stores calibration data in .ini format.
5.3 Experimental Calibration of the Alignment Platform
Using three-point calibration data to guide alignment resulted in large alignment errors. Analysis attributed errors to differences between ideal model and actual model. Manufacturing and assembly errors caused motor installation positions to deviate from ideal design positions, so motor pulse values cannot directly represent platform coordinate positions. Improving machining and assembly precision has limited capability. External measurement devices such as laser interferometers are expensive, and error sources are numerous, making individual measurement and compensation difficult. Therefore, experimental calibration was ultimately adopted.
After four-camera calibration, deviations $\Delta x$, $\Delta y$, $\Delta\theta$ are obtained in the upper-left camera coordinate system. The system must derive pulse increments $\Delta u$, $\Delta v$, $\Delta w$ for the three motors. This is a three-input, three-output model. The inputs and outputs are nonlinear, but nonlinear theory is insufficiently developed. Engineering projects with dead zones, gaps, and other nonlinear phenomena typically linearize nonlinear systems. This approach is also adopted here, approximating the nonlinear relationship as linear:
$$\begin{bmatrix} \Delta u \\ \Delta v \\ \Delta w \end{bmatrix} = \begin{bmatrix} w_{11} & w_{12} & w_{13} \\ w_{21} & w_{22} & w_{23} \\ w_{31} & w_{32} & w_{33} \end{bmatrix} \begin{bmatrix} \Delta x \\ \Delta y \\ \Delta \theta \end{bmatrix} + \begin{bmatrix} m_1 \\ m_2 \\ m_3 \end{bmatrix}$$
There are 12 coefficients to solve, requiring 12 equations. Each input-output group provides 3 equations, so 4 groups of input-output data are needed. Experiments were conducted multiple times: panels were placed in position, U, V, W motors were moved multiple times with moderate distances ensuring Mark points remained in camera fields of view, motor movements were recorded, and deviations in the upper-left camera coordinate system were calculated using calibration relationships. Results showed coefficients fluctuated within certain ranges, so averaging multiple calculation results was adopted. Error analysis on 5 additional input-output groups showed the W motor calculation error does not exceed 10 pulses; U and V motor errors are similar, not exceeding 40 pulses. With motor resolution of 400 pulses per millimeter, 0.1 mm error is below the required 0.15 mm alignment accuracy, so the linear model is acceptable.
Chapter 6 System Operation and Analysis
After completing platform calibration, the linear relationship between panel deviation ($\Delta x$, $\Delta y$, $\Delta\theta$) and motor pulse increments ($\Delta u$, $\Delta v$, $\Delta w$) was established. Calculated pulse increments are decimals requiring rounding to integers. Since the linear model approximates the actual nonlinear relationship, a single correction cycle may not meet precision requirements. Therefore, closed-loop control is introduced: multiple alignment cycles are performed. After each cycle, panel deviations are recalculated; if requirements are not met, alignment continues until precision is satisfied.
All algorithms and calibration data were integrated into the software system. When panels are placed by robots, Mark points are located first. Coordinate information from each camera image is converted to the upper-left image coordinate system using four-camera calibration data. Deviations $\Delta x$, $\Delta y$, $\Delta\theta$ are calculated, the linear relationship model drives U, V, W motors, and motor positions are verified before recalculating deviations. Typically, two alignment cycles are sufficient.
To verify system accuracy, 10 finished CPV modules were randomly selected and measured using a 3D CNC measuring instrument.
Table 6-1 Position Deviations of Panels After Alignment
| No. | X Deviation [mm] | Y Deviation [mm] | Angle Deviation [°] |
|---|---|---|---|
| 1 | 0.0731 | 0.0823 | 0.0095 |
| 2 | 0.0716 | 0.0815 | 0.0139 |
| 3 | 0.0978 | 0.0642 | 0.0114 |
| 4 | 0.0791 | 0.0725 | 0.0131 |
| 5 | 0.0774 | 0.0893 | 0.0144 |
| 6 | 0.0815 | 0.0796 | 0.0104 |
| 7 | 0.0917 | 0.0634 | 0.0135 |
| 8 | 0.0818 | 0.0780 | 0.0124 |
| 9 | 0.0720 | 0.0829 | 0.0123 |
| 10 | 0.0637 | 0.0821 | 0.0118 |
| Average | 0.0790 | 0.0776 | 0.0126 |
| Std Dev | 0.0100 | 0.0085 | 0.0012 |
| Range | 0.0341 | 0.0259 | 0.0040 |
Analysis of experimental data shows position deviations of panels after alignment using this system do not exceed 0.12 mm, with standard deviation not exceeding 0.02 mm, meeting project precision requirements. The system has entered the trial production stage of CPV modules for further testing and audit.
Conclusion and Outlook
To achieve precise alignment between upper and lower photovoltaic panels of CPV modules and ensure photoelectric conversion efficiency, this thesis researched and designed a visual alignment system for high-concentrating photovoltaic panels. The system can conveniently perceive panel positions and postures through machine vision and guide the alignment platform to achieve accurate alignment based on calibration algorithms. The main work and innovations are as follows:
(1) Based on Mark point position distribution and characteristics, a multi-camera array vision scheme was designed, with four cameras acquiring Mark point images. The entire visual alignment system was modularized into motion control and machine vision components. Based on precision requirements, industrial cameras were selected; telecentric lenses were chosen to eliminate perspective distortion; after multiple attempts, a red high-power ring light illumination scheme was established for both upper and lower camera groups.
(2) Current circular Mark point positioning algorithms were researched and compared. Considering actual Mark point image conditions — adhesive interference and missing points — the shape-based template matching algorithm was determined. Through image processing algorithms, Mark point geometric features (circular contours) were extracted. The similarity measure function was determined, and image pyramids accelerated matching. The algorithm was evaluated for stability and runtime.
(3) Planar coordinate system calibration principles were introduced. After analyzing calibration benchmark requirements, a calibration standard was designed for this system’s four-camera calibration. Calibration methods between each camera and the upper-left camera were developed. Software was written to complete four-camera image coordinate system calibration.
(4) Kinematic modeling of the alignment platform was established, and the three-point calibration method was introduced. A practical implementation using the lower panel and existing calibration relationships was designed. Calibration software was written. When precision requirements were not satisfied, experimental calibration was adopted, solving the input-output linearized model through experimental data. All algorithms and calibration data were integrated into the final system software, achieving the required alignment precision.
Some aspects deserve further research:
(1) Although telecentric lenses eliminate perspective distortion, extremely small errors still exist in each camera image. Individual camera calibration could eliminate these errors.
(2) The calibration standard required high manufacturing precision, making it costly. Better calibration methods could be researched.
(3) The alignment platform’s manufacturing and assembly precision limitations caused three-point calibration results to be less than ideal. The nonlinear relationship was approximated as linear through experimental modeling. Intelligent algorithms such as artificial neural networks could provide better results.
